VLDB 2026 Research / reviewers in the wild / expert
Xiaodan Hu
dblp:09/9830
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional diffusion denoising for robust social recommendation with contrastive learning and curriculum learning strategies
Shun Sun, Xiaodan Hu, Fuzhen Sun |
Data Knowl. Eng. | 4 |
| 2026 | Mask Balancing: Perception-Driven Dynamic Visibility Enhancement for Occlusion-Capable Optical See-Through Head-Mounted DisplaysabstractThe poor transparency of occlusion-capable optical see-through head-mounted displays (OC-OSTHMDs) deteriorates the visibility of the real scene, hindering the practical application of the devices. Previous works mitigate the issue by upgrading the transmittance of the spatial light modulator (SLM). However, the strategy soon reaches a limit because further optimization requires improving the transmittance of all optical elements, e.g., lenses and beam splitters. Moreover, pixelated occlusion usually relies on polarizing the real scene light, inevitably cutting the input optical power by half. To overcome this limitation, we propose a mask balancing method that improves real-scene brightness through polarization blending. Specifically, the s-polarized component, which passes through the optical system to provide occlusion-capable vision, is blended with the p-polarized component, which bypasses the system to preserve the raw view of the real scene. The blending is realized by simply modulating the cross-angle between a polarizing beam splitter and a linear polarizer, benefiting the robustness and versatility of the proposed method. We introduce a perception-driven blending approach, where the cross-angle is optimized in real-time to balance the visibility of the real scene and the texture and lighting of the virtual object. A benchtop prototype is built. A user study with 12 participants is conducted to quantify the visibility threshold of the texture and lighting of virtual objects. Then, a user study with 12 participants proves that the proposed method improves the visibility of the real scene while keeping a good appearance of the virtual object. We believe the proposed method is an important step toward developing practical solutions for OC-OSTHMDs. Yan Zhang 0101, Rundong Chu, Qingtai Dong, Xiaodan Hu, Keyao You, Zixuan Guo 0003, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | X-Mask: Improving Soft-Edge Occlusion in Optical See-Through Displays with Cross-Shaped PinholesabstractPlacing a transparent liquid crystal display (LCD) into the light path is a simple approach to create occlusion-capable optical seethrough head-mounted displays (OST-HMDs) that suffers from defocused (soft-edge) occlusion where the mask leakage partially occludes surrounding content as well. Creating a focused (hard-edge) occlusion that does not suffer from mask leakage requires complicated, bulky optical setups. We present X-Mask, a pinhole-arraybased OST-HMD that creates a sharp occlusion mask without the need for a bulky setup requiring only two transparent LCD layers. By rendering a pinhole array on the layer closer to the user's eye, our system functions as a programmable aperture layer that extends the effective depth of field and improves the sharpness of the occlusion mask rendered on the second LCD layer. Utilizing a conventional circular pinhole would result in non-uniform brightness and contrast. By changing the pinhole shape to a cross enables nearoptimal retinal tiling with reduced overlaps and gaps. To accommodate pupil size variation, focus distance, and gaze direction, our system design allows for gaze-contingent adjustment of both LCD layers. We validate X-Mask in simulations and a physical prototype showing improved occlusion sharpness and visual uniformity. Xiaodan Hu, Christoph Ebner, Yan Zhang 0101, Kiyoshi Kiyokawa, Alexander Plopski |
ISMAR | 1 |
| 2025 | Perception-Driven Soft-Edge Occlusion for Optical See-Through Head-Mounted DisplaysabstractSystems with occlusion capabilities, such as those used in vision augmentation, image processing, and optical see-through head-mounted display (OST-HMD), have gained popularity. Achieving precise (hard-edge) occlusion in these systems is challenging, often requiring complex optical designs and bulky volumes. On the other hand, utilizing a single transparent liquid crystal display (LCD) is a simple approach to create occlusion masks. However, the generated mask will appear defocused (soft-edge) resulting in insufficient blocking or occlusion leakage. In our work, we delve into the perception of soft-edge occlusion by the human visual system and present a preference-based optimal expansion method that minimizes perceived occlusion leakage. In a user study involving 20 participants, we made a noteworthy observation that the human eye perceives a sharper edge blur of the occlusion mask when individuals see through it and gaze at a far distance, in contrast to the camera system's observation. Moreover, our study revealed significant individual differences in the perception of soft-edge masks in human vision when focusing. These differences may lead to varying degrees of demand for mask size among individuals. Our evaluation demonstrates that our method successfully accounts for individual differences and achieves optimal masking effects at arbitrary distances and pupil sizes. Xiaodan Hu, Yan Zhang 0101, Alexander Plopski, Yuta Itoh 0001, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | ReAR Indicators: Peripheral Cycling Indicators for Rear-Approaching HazardsabstractDuring cycling activities, cyclists often focus on pedestrians, vehicles or road conditions in front of their bicycle. Because of this forward focus, approaching vehicles from behind can easily be missed, which can result in accidents, injury, or death. Although rear information can viewed with handle-mounted mirrors or monitors, looking down can distract the cyclist from other hazards. Guanghan Zhao, Xiaodan Hu, Jason Orlosky, Kiyoshi Kiyokawa |
AVI | 2 |
| 2024 | Retinotopic Foveated RenderingabstractFoveated rendering (FR) improves the rendering performance of virtual reality (VR) by allocating fewer computational loads in the peripheral field of view (FOV). Existing FR techniques are built based on the radially symmetric regression model of human visual acuity. However, horizontal-vertical asymmetry (HVA) and vertical meridian asymmetry (VMA) in the cortical magnification factor (CMF) of the human visual system have been evidenced by retinotopy research of neuroscience, suggesting the radially asymmetric regression of visual acuity. In this paper, we begin with functional magnetic resonance imaging (fMRI) data, construct an anisotropic CMF model of the human visual system, and then introduce the first radially asymmetric regression model of the rendering precision for FR applications. We conducted a pilot experiment to adapt the proposed model to VR head-mounted displays (HMDs). A user study demonstrates that retinotopic foveated rendering (RFR) provides participants with perceptually equal image quality compared to typical FR methods while reducing fragments shading by 27.2% averagely, leading to the acceleration of 1/6 for graphics rendering. We anticipate that our study will enhance the rendering performance of VR by bridging the gap between retinotopy research in neuroscience and computer graphics in VR. Yan Zhang 0101, Keyao You, Xiaodan Hu, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang |
VR | 3 |
| 2023 | Add-on Occlusion: Turning Off-the-Shelf Optical See-through Head-mounted Displays Occlusion-capableabstractThe occlusion-capable optical see-through head-mounted display (OC-OSTHMD) is actively developed in recent years since it allows mutual occlusion between virtual objects and the physical world to be correctly presented in augmented reality (AR). However, implementing occlusion with the special type of OSTHMDs prevents the appealing feature from the wide application. In this paper, a novel approach for realizing mutual occlusion for common OSTHMDs is proposed. A wearable device with per-pixel occlusion capability is designed. OSTHMD devices are upgraded to be occlusion-capable by attaching the device before optical combiners. A prototype with HoloLens 1 is built. The virtual display with mutual occlusion is demonstrated in real-time. A color correction algorithm is proposed to mitigate the color aberration caused by the occlusion device. Potential applications, including the texture replacement of real objects and the more realistic semi-transparent objects display, are demonstrated. The proposed system is expected to realize a universal implementation of mutual occlusion in AR. Yan Zhang 0101, Xiaodan Hu, Kiyoshi Kiyokawa, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Unsupervised 3D Pose Estimation for Hierarchical Dance Video Recognition *abstractDance experts often view dance as a hierarchy of information, spanning low-level (raw images, image sequences), mid-levels (human poses and bodypart movements), and high-level (dance genre). We propose a Hierarchical Dance Video Recognition framework (HDVR). HDVR estimates 2D pose sequences, tracks dancers, and then simultaneously estimates corresponding 3D poses and 3D-to-2D imaging parameters, without requiring ground truth for 3D poses. Unlike most methods that work on a single person, our tracking works on multiple dancers, under occlusions. From the estimated 3D pose sequence, HDVR extracts body part movements, and therefrom dance genre. The resulting hierarchical dance representation is explainable to experts. To overcome noise and interframe correspondence ambiguities, we enforce spatial and temporal motion smoothness and photometric continuity over time. We use an LSTM network to extract 3D movement subsequences from which we recognize dance genre. For experiments, we have identified 154 movement types, of 16 body parts, and assembled a new University of Illinois Dance (UID) Dataset, containing 1143 video clips of 9 genres covering 30 hours, annotated with movement and genre labels. Our experimental results demonstrate that our algorithms outperform the state-of-the-art 3D pose estimation methods, which also enhances our dance recognition performance. Xiaodan Hu, Narendra Ahuja |
ICCV | 1 |
| 2021 | Non-stationary content-adaptive projector resolution enhancement
Xiaodan Hu, Mohamed A. Naiel, Zohreh Azimifar, Ibrahim Ben Daya, Mark Lamm, Paul W. Fieguth |
Signal Process. Image Commun. | 1 |
| 2020 | Squeeze-and-Attention Networks for Semantic SegmentationabstractThe recent integration of attention mechanisms into segmentation networks improves their representational capabilities through a great emphasis on more informative features. However, these attention mechanisms ignore an implicit sub-task of semantic segmentation and are constrained by the grid structure of convolution kernels. In this paper, we propose a novel squeeze-and-attention network (SANet) architecture that leverages an effective squeeze-and-attention (SA) module to account for two distinctive characteristics of segmentation: i) pixel-group attention, and ii) pixel-wise prediction. Specifically, the proposed SA modules impose pixel-group attention on conventional convolution by introducing an 'attention' convolutional channel, thus taking into account spatial-channel inter-dependencies in an efficient manner. The final segmentation results are produced by merging outputs from four hierarchical stages of a SANet to integrate multi-scale contexts for obtaining an enhanced pixel-wise prediction. Empirical experiments on two challenging public datasets validate the effectiveness of the proposed SANets, which achieves 83.2 % mIoU (without COCO pre-training) on PASCAL VOC and a state-of-the-art mIoU of 54.4 % on PASCAL Context. Zilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu, Ibrahim Ben Daya, Wei-Shi Zheng 0001, Jonathan Li 0001, Alexander Wong |
CVPR | 4 |
| 2011 | Recovery from Link Failures in Networks with Arbitrary Topology via Diversity CodingabstractLink failures in wide area networks are common. To recover from such failures, a number of methods such as SONET rings, protection cycles, and source rerouting have been investigated. Two important considerations in such approaches are the recovery time and the needed spare capacity to complete the recovery. Usually, these techniques attempt to achieve a recovery time less than 50 ms. In this paper we introduce an approach that provides link failure recovery in a hitless manner, or without any appreciable delay. This is achieved by means of a method called diversity coding. We present an algorithm for the design of an overlay network to achieve recovery from single link failures in arbitrary networks via diversity coding. This algorithm is designed to minimize spare capacity for recovery. We compare the recovery time and spare capacity performance of this algorithm against conventional techniques in terms of recovery time, spare capacity, and a joint metric called Quality of Recovery (QoR). QoR incorporates both the spare capacity percentages and worst case recovery times. Based on these results, we conclude that the proposed technique provides much shorter recovery times while achieving similar extra capacity, or better QoR performance overall. Serhat Nazim Avci, Xiaodan Hu, Ender Ayanoglu |
GLOBECOM | 2 |