VLDB 2026 Research / reviewers in the wild / expert
Yiren Zhou
dblp:55/6521
· DBLP profile ↗
13ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Single-view Image to Novel-view Generation for Hand-Object InteractionsabstractHand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view hand-object interaction images from a single image. To this end, we first train a 2D diffusion prior. Given the camera pose in novel views, our approach transfers the camera information into explicit hand representations, including hand depth and skeleton images. We propose a global hand embedding to control the diffusion model based on these hand representations. We then learn a 3D Gaussian splatting for novel-view rendering using the diffusion prior. However, occluded objects present a persistent challenge. To address this issue, we further introduce local hand embedding, where a contact field is defined in the 3D Gaussian Splatting. We leverage contact information to guide the rendering in the contact field. Extensive experiments on the HO3D and DexYCB datasets demonstrate that our method significantly outperforms state-of-the-art novel-view synthesis for hand-object interactions. Zhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, Jifei Song |
AAAI | 4 |
| 2025 | Unraveling the Characterization and Propagation of Security Vulnerabilities in TensorFlow-based Deep Learning Software Supply ChainabstractAs a widely adopted deep learning (DL) framework, TensorFlow's vulnerabilities have the potential to affect a substantial number of DL applications throughout the software supply chain (SSC).However, Existing research lacks a comprehensive exploration of the characterization of TensorFlow vulnerabilities and the propagation of vulnerabilities on SSC.To help TensorFlow-based developers and security specialists in understanding the security risks, we construct an empirical study on 429 vulnerabilities of TensorFlow across a TensorFlow-based vulnerability SSC comprising 5790 versions across 691 affected packages constructed through the GitHub dependency graph.We observe that: 1) A predominant share (79.6%) of vulnerabilities occur in the TensorFlow's Core modules (e.g., kernels, and ops), featuring prevalent vulnerabilities such as Reachable Assertion, Out-of-bounds Read, Improper Input Validation, and NULL Pointer Dereference.Notably, Divide By Zero vulnerabilities pose significant risks in the Lite module; 2) Vulnerability co-occurrence is observed in 20 pairs involving 50 vulnerabilities, with Heap-based Buffer Overflow vulnerabilities particularly likely to coexist with other types of vulnerabilities; 3) Packages within the Large Language Models (LLM) domain, often distributed across the third and fourth layers of the SSC, are vulnerable to TensorFlow's security issues; 4) Many commonly utilized APIs (i.e.tf.constant, tf.concat, and tf.range), are implicated in TensorFlow vulnerabilities, affecting over half of all packages and, by extension, a significant portion of software within the SSC.Our findings suggest that: i)It would be better for Tensorflow developers take input validation, bounds checking, and assertion and exception handling before performing tensor operations, division operations, and pointer accesses to avoid Reachable Assertion, Divide By Zero, and memory-related vulnerabilities.ii) TensorFlow developers would be better to scrutinize the presence of memory-related vulnerabilities when encountering a vulnerability stemming from improper input validation.iii)TensorFlow-based Developers would be better to be aware of * Yiren Zhou, Lina Gong |
Internetware | 1 |
| 2024 | NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object InteractionabstractModeling hand-object interactions is a fundamentally challenging task in 3D computer vision. Despite remarkable progress that has been achieved in this field, existing methods still fail to synthesize the hand-object interaction photo-realistically, suffering from degraded rendering quality caused by the heavy mutual occlusions between the hand and the object, and inaccurate hand-object pose estimation. To tackle these challenges, we present a novel free-viewpoint rendering framework, Neural Contact Radiance Field (NCRF), to reconstruct hand-object interactions from a sparse set of videos. In particular, the proposed NCRF framework consists of two key components: (a) A contact optimization field that predicts an accurate contact field from 3D query points for achieving desirable contact between the hand and the object. (b) A hand-object neural radiance field to learn an implicit hand-object representation in a static canonical space, in concert with the specifically designed hand-object motion field to produce observation-to-canonical correspondences. We jointly learn these key components where they mutually help and regularize each other with visual and geometric constraints, producing a high-quality hand-object reconstruction that achieves photorealistic novel view synthesis. Extensive experiments on HO3D and DexYCB datasets show that our approach outperforms the current state-of-the-art in terms of both rendering quality and pose estimation accuracy. Zhongqun Zhang, Jifei Song, Eduardo Pérez-Pellitero, Yiren Zhou, Hyung Jin Chang, Ales Leonardis |
3DV | 4 |
| 2024 | Human Gaussian Splatting: Real-Time Rendering of Animatable AvatarsabstractThis work addresses the problem of real-time rendering of photorealistic human body avatars learned from multi-view videos. While the classical approaches to model and render virtual humans generally use a textured mesh, recent research has developed neural body representations that achieve impressive visual quality. However, these models are difficult to render in real-time and their quality degrades when the character is animated with body poses different than the training observations. We propose an animatable human model based on 3D Gaussian Splatting, that has recently emerged as a very efficient alternative to neural radiance fields. The body is represented by a set of gaussian primitives in a canonical space which is deformed with a coarse to fine approach that combines forward skinning and local non-rigid refinement. We describe how to learn our Human Gaussian Splatting (HuGS) model in an end-to-end fashion from multi-view observations, and evaluate it against the state-of-the-art approaches for novel pose synthesis of clothed body. Our method achieves 1.5 dB PSNR improvement over the state-of-the-art on THuman4 dataset while being able to render in real-time (≈ 80 fps for$512 \times 512$resolution). Arthur Moreau, Jifei Song, Helisa Dhamo, Richard Shaw, Yiren Zhou, Eduardo Pérez-Pellitero |
CVPR | 5 |
| 2024 | HeadGaS: Real-Time Animatable Head Avatars via 3D Gaussian Splatting
Helisa Dhamo, Yinyu Nie, Arthur Moreau, Jifei Song, Richard Shaw, Yiren Zhou, Eduardo Pérez-Pellitero |
ECCV (2) | 6 |
| 2024 | Mitigating Intra-host Network Congestion with SmartNICabstractWith the rapid development and wide deployment of high-speed network technologies like RDMA and the relatively stagnant evolution of intra-host resources, intra-host network congestion has become a potential issue that may affect the QoS of network applications. Offloading hotspot data to modern Smart-NICs, enabling hotspot data access completion on the SmartNIC, and reducing intra-host network traffic, is a promising solution to this issue. However, due to the limited SmartNIC resources and the complexity of network application requirements, achieving efficient offload is challenging.We present Magician, an architecture to mitigate intra-host network congestion with SmartNIC. Magician adopts a client-driven data access approach to avoid performance degradation caused by limited SmartNIC resources. Magician also introduces a SmartNIC-oriented hotspot data update strategy that dynamically refreshes hotspot data with minimal overhead. Moreover, we design a server-centric data consistency mechanism to ensure data consistency under concurrent access. We implement Magician within the key-value store. Evaluation of the key-value store with and without Magician suggests that, in the presence of intra-host network congestion, Magician significantly mitigates intra-host network congestion, leading to improved performance of network applications. Lulu Chen, Chunpu Huang, Rui Zhang 0112, Yiren Zhou, Ming Yan 0009, Jie Wu 0003 |
IWQoS | 5 |
| 2024 | TCAMVisor: High-throughput TCAM Virtualization for Multi-tenant Software Defined NetworkingabstractSoftware Defined Networking (SDN) provides users with a unified abstraction of physical networks. To meet the demands of modern data centers, many works have focused on designing network virtualization hypervisors that support multi-tenant SDN. Ternary Content Addressable Memory (TCAM) is widely used in SDN switches for rule storage. While it has extremely high lookup throughput, it also features drawbacks such as small capacity and slow update speed. Faced with multi-tenant scenarios, its limitations are even more pronounced. Existing hypervisors lack consideration for TCAM isolation, leading to slower TCAM updates and the mutual impact of requests from different tenants. Consequently, they fail to provide guaranteed performance to tenants. To solve these problems, we propose TCAMVisor, which further isolates TCAM resources based on traditional SDN hypervisors. Specifically, TCAMVisor provides better allocation mechanisms for TCAM entry and control bandwidth, ensuring inter-tenant isolation while improving resource utilization. Additionally, TCAMVisor improves the update speed of TCAM by delicately placing tenant rules. To the best of our knowledge, TCAMVisor is the first work to effectively achieve tenant isolation in TCAM, with an average throughput improvement of 5.5 times compared to FlowVisor. Ruoshi Sun, Ruyi Yao, Hao Wang 0231, Yiren Zhou, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 5 |
| 2023 | On the Importance of Accurate Geometry Data for Dense 3D Vision TasksabstractLearning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less regions are problematic for structure from motion and stereo, reflective material poses issues for active sensing, and distances for translucent objects are intricate to measure with existing hardware. Training on inaccurate or corrupt data induces model bias and hampers generalisation capabilities. These effects remain unnoticed if the sensor measurement is considered as ground truth during the evaluation. This paper investigates the effect of sensor errors for the dense 3D vision tasks of depth estimation and reconstruction. We rigorously show the significant impact of sensor characteristics on the learned predictions and notice generalisation issues arising from various technologies in everyday household environments. For evaluation, we introduce a carefully designed dataset11dataset available at https://github.com/Junggy/HAMMER-dataset comprising measurements from commodity sensors, namely D-ToF, I-ToF, passive/active stereo, and monocular RGB+P. Our study quantifies the considerable sensor noise impact and paves the way to improved dense vision estimates and targeted data fusion. Patrick Ruhkamp, Guangyao Zhai, Nikolas Brasch, Yannick Verdie, Jifei Song, Yiren Zhou, Anil Armagan, Slobodan Ilic, Ales Leonardis, Nassir Navab, Benjamin Busam |
CVPR | 8 |
| 2022 | Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment
Xinghui Li, Benjamin Busam, Yiren Zhou, Ales Leonardis, Shanxin Yuan |
BMVC | 4 |
| 2018 | Adaptive Quantization for Deep Neural NetworkabstractIn recent years Deep Neural Networks (DNNs) have been rapidly developed in various applications, together with increasingly complex architectures. The performance gain of these DNNs generally comes with high computational costs and large memory consumption, which may not be affordable for mobile platforms. Deep model quantization can be used for reducing the computation and memory costs of DNNs, and deploying complex DNNs on mobile equipment. In this work, we propose an optimization framework for deep model quantization. First, we propose a measurement to estimate the effect of parameter quantization errors in individual layers on the overall model prediction accuracy. Then, we propose an optimization process based on this measurement for finding optimal quantization bit-width for each layer. This is the first work that theoretically analyse the relationship between parameter quantization errors of individual layers and model accuracy. Our new quantization algorithm outperforms previous quantization optimization methods, and achieves 20-40% higher compression rate compared to equal bit-width quantization at the same model prediction accuracy. Yiren Zhou, Seyed-Mohsen Moosavi-Dezfooli, Ngai-Man Cheung, Pascal Frossard |
AAAI | 1 |
| 2018 | Computation and Memory Efficient Image SegmentationabstractIn this paper, we address the segmentation problem under limited computation and memory resources. Given a segmentation algorithm, we propose a framework that can reduce its computation time and memory requirementsimultaneously, while preserving its accuracy. The proposed framework uses standard pixel-domain downsampling and includes two main steps.Coarse segmentationis first performed on the downsampled image.Refinementis then applied to the coarse segmentation results. We make two novel contributions to enable competitive accuracy using this simple framework. First, we rigorously examine the effect of downsampling on segmentation using a signal processing analysis. The analysis helps to determine theuncertain regions, which are small image regions where pixel labels are uncertain after the coarse segmentation. Second, we propose an efficient minimum spanning tree-based algorithm to propagate the labels into the uncertain regions. We perform extensive experiments using several standard data sets. The experimental results show that our segmentation accuracy is comparable to state-of-the-art methods, while requiring much less computation time and memory than those methods. Yiren Zhou, Thanh-Toan Do, Haitian Zheng, Ngai-Man Cheung, Lu Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Accessible Melanoma Detection Using Smartphones and Mobile Image AnalysisabstractWe investigate the design of an entire mobile imaging system for early detection of melanoma. Different from previous work, we focus on smartphone-captured visible light images. Our design addresses two major challenges. First, images acquired using a smartphone under loosely-controlled environmental conditions may be subject to various distortions, and this makes melanoma detection more difficult. Second, processing performed on a smartphone is subject to stringent computation and memory constraints. In our work, we propose a detection system that is optimized to run entirely on the resource-constrained smartphone. Our system intends to localize the skin lesion by combining a lightweight method for skin detection with a hierarchical segmentation approach using two fast segmentation methods. Moreover, we study an extensive set of image features and propose new numerical features to characterize a skin lesion. Furthermore, we propose an improved feature selection algorithm to determine a small set of discriminative features used by the final lightweight system. In addition, we study the human-computer interface (HCI) design to understand the usability and acceptance issues of the proposed system. Our extensive evaluation on an image dataset provided by National Skin Center - Singapore (117 benign nevi and 67 malignant melanoma) confirms the effectiveness of the proposed system for melanoma detection: 89.09% sensitivity at specificity ≥90%. Thanh-Toan Do, Tuan Hoang, Victor Pomponiu, Yiren Zhou, Zhao Chen 0005, Ngai-Man Cheung, Dawn Chin-Ing Koh, Aaron Tan, Suat-Hoon Tan |
IEEE Trans. Multim. | 4 |
| 2017 | On classification of distorted images with deep convolutional neural networksabstractImage blur and image noise are common distortions during image acquisition. In this paper, we systematically study the effect of image distortions on the deep neural network (DNN) image classifiers. First, we examine the DNN classifier performance under four types of distortions. Second, we propose two approaches to alleviate the effect of image distortion: re-training and fine-tuning with noisy images. Our results suggest that, under certain conditions, fine-tuning with noisy images can alleviate much effect due to distorted inputs, and is more practical than re-training. Yiren Zhou, Sibo Song, Ngai-Man Cheung |
ICASSP | 1 |