EDBT 2026 Demo / reviewers in the wild / expert
Xiaobai Chen
dblp:65/7188
· DBLP profile ↗
22ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0003-1166-6927ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FineRR-ZNS: Enabling Fine-Granularity Read Refreshing for ZNS SSDsabstractZoned namespace (ZNS) SSDs are emerging storage devices offering low cost, high performance, and software definability. By adopting host-managed zone-based sequential programming, ZNS SSDs effectively eliminate the space overhead associated with on-board DRAM memory and garbage collection. However, while background read refreshing serves as a data protection mechanism in conventional block-interface SSDs, state-of-the-art ZNS SSDs lack read refreshing functionality to guarantee data reliability. Moreover, implementing zonelevel read refreshing in ZNS SSDs incurs significant overhead due to the large volume of valid data movements in a zone, leading to degraded I/O performance. To efficiently enable read refreshing for ZNS SSDs, this paper proposes FineRR-ZNS, a fine-granularity read refreshing mechanism for ZNS SSDs. FineRR-ZNS employs a host-controlled fine-granularity read refreshing scheme that selectively determines block-level read refreshing via metadata remapping. A zone reconstruction method is also designed to retrieve remapped data forming complete data during zone-level RR. Specifically, the remapped data after zone reconstruction are still available and prioritized for read access until their respective blocks need the next RR. Evaluation results show that FineRR-ZNS significantly enhances read refreshing efficiency and I/O throughput compared to zone-level read refreshing implemented in the state-of-the-art ZenFS file system. Jun Li 0062, Zhibing Sha, Fan Yang 0110, Xiaofei Xu 0002, Xiaobai Chen, Jieming Yin, Jianwei Liao 0001 |
DAC | 5 |
| 2025 | CoupledCB: Eliminating Wasted Pages in Copyback-based Garbage Collection for SSDsabstractThe management of garbage collection poses significant challenges in high-density NAND flash-based SSDs. The introduction of the copyback command aims to expedite the migration of valid data. However, its odd/even constraint causes wasted pages during migrations, limiting the efficiency of garbage collection. Additionally, while full-sequence programming en-hances write performance in high-density SSDs, it increases write granularity and exacerbates the issue of wasted pages. To address the problem of wasted pages, we propose a novel method called CoupledCB, which utilizes coupled blocks to fill up the wasted space in copyback-based garbage collection. By taking into account the access characteristics of the candidate coupled blocks and workloads, we develop a coupled block selection model assisted by logistic regression. Experimental results show that our proposal significantly enhances garbage collection efficiency and 1/O performance compared to state-of-the-art schemes. Jun Li 0062, Xiaofei Xu 0002, Zhibing Sha, Xiaobai Chen, Jieming Yin, Jianwei Liao 0001 |
DATE | 4 |
| 2025 | DriveGPT: Scaling Autoregressive Behavior Models for DrivingabstractWe present DriveGPT, a scalable behavior model for autonomous driving. We model driving as a sequential decision-making task, and learn a transformer model to predict future agent states as tokens in an autoregressive fashion. We scale up our model parameters and training data by multiple orders of magnitude, enabling us to explore the scaling properties in terms of dataset size, model parameters, and compute. We evaluate DriveGPT across different scales in a planning task, through both quantitative metrics and qualitative examples, including closed-loop driving in complex real-world scenarios. In a separate prediction task, DriveGPT outperforms state-of-the-art baselines and exhibits improved performance by pretraining on a large-scale dataset, further validating the benefits of data scaling. Eric M. Wolff, Paul Vernaza, Tung Phan-Minh, Hongge Chen, David S. Hayden, Mark Edmonds, Brian Pierce, Xinxin Chen, Pratik Elias Jacob, Xiaobai Chen, Chingiz Tairbekov, Pratik Agarwal, Tianshi Gao, Yuning Chai, Siddhartha S. Srinivasa |
ICML | 11 |
| 2025 | CAMC: A Multi-Chiplet Accelerator With Heterogeneous Memory-Based Computing Architecture For DNN TrainingabstractDeep Neural Networks (DNNs) are extensively utilized in various fields due to their remarkable performance. However, as DNN models increase in complexity and size, the training process incurs substantial data transfer costs between computation and storage. The slowdown of Moore’s Law further challenges the integration of additional resources on a single chip, making it difficult to improve storage capacity and reduce off-chip data transfers. To address these challenges, we propose CAMC, a multi-chiplet DNN accelerator with a heterogeneous memory computing architecture. CAMC integrates SRAM-based in-memory computing with TSV-stacked DRAM-based near-memory computing. In addition, an efficient mapping strategy was developed to optimize resource utilization and performance. The experimental results demonstrate that CAMC enhances energy efficiency by 11.84 times and reduces data transfer costs by 8.78 times compared to the baseline design. Xiaobai Chen, Jiacheng Mei, Yifei Tian, Jieming Yin, Fu Xiao 0001 |
ISCAS | 1 |
| 2025 | DPAcc: An FPGA-based Differential Privacy Acceleration FrameworkabstractIn the data-driven era, privacy protection has become a critical concern. Differential privacy is an effective technique that incorporates random noise during data processing to ensure that alterations to individual data points do not significantly affect overall outputs. However, the additional operations required by differential privacy can result in prolonged training times and degraded model performance. This work proposes DPAcc, an FPGA-based acceleration framework for differential privacy that utilizes hardware implementation to decrease training time. The designed FPGA module efficiently executes clipping and noise addition operations, significantly reducing the overhead compared to standard training. Experimental results demonstrate that DPAcc improves training efficiency across multiple models, achieving up to 2× speedup compared to standard differential privacy training methods. Ao Dong, Pengyang Li, Yifei Tian, Xiaobai Chen, Jieming Yin |
ISCAS | 5 |
| 2025 | FlexAcc: Accelerating Batch Normalization through GPU-FPGA IntegrationabstractConvolutional neural networks are fundamental to deep learning, especially in computer vision. However, their computational demands, particularly during batch normalization, create significant inefficiencies due to excessive data movement between memory and processing units. To address this, we propose FlexAcc, a novel architecture that integrates GPUs and FPGAs to offload BN computations to FPGAs, reducing data movement and improving hardware utilization. FlexAcc accelerates end-to-end training performance by up to 1.1× across various models. This approach bridges the performance gap between convolutional and non-convolutional layers, advancing deep learning model deployment. Haishuai Zhang, Pengyang Li, Xuehuai Shi, Xiaobai Chen, Jieming Yin |
ISCAS | 5 |
| 2025 | MMG: Manipulation-Aware Holistic Human Motion Generation from Sparse Tracking SignalsabstractGenerating realistic avatar motion via sparse tracking signals through VR devices is essential for enhancing the immersive user experience. Human-object manipulation behaviors not only affect hand motion but also significantly impact body motion. However, existing motion generation methods for human-object interactions overlook the coordinated coupling between body and hand motions during manipulations. Due to the diversity and complexity of holistic motion (body and hand motions simultaneously) in the latent motion space, generating physically plausible and temporally consistent holistic motion in real time, via the joint constraints imposed by sparse tracking signals and manipulation content, is a major challenge in the human motion generation task. We propose the manipulation-aware holistic human motion generation method (MMG) to help resolve this issue. In MMG, first, we construct a manipulation-aware holistic human motion generation framework that serially compresses the latent motion space distribution of the body and hand to generate realistic holistic human motion with object manipulation enabled. Second, to enhance the impact of object manipulation on holistic motion generation, MMG designs a novel object manipulation representation to extract effective manipulation features. Third, MMG is trained by an elaborate progressive manipulation-guided training algorithm to improve motion generation robustness and inference performance. Compared to state-of-the-art methods, MMG achieves up to a 39% improvement in the generated holistic motion quality with a 3.55 × speedup in generation performance. In manipulation-enabled scenes, MMG generates holistic motion in real time ($\geq 24 f p s$). Compared to the state-of-the-art methods, its perceived quality is significantly improved, and the task performance of holistic motion-required VR manipulation is high-significantly improved. This paper's code is at https://github.com/XRZ-BUAA/MMG. Xuehuai Shi, Renzhi Xiao, Yilun Sheng, Xiaobai Chen, Jieming Yin, Qingshan Liu 0001 |
ISMAR | 6 |
| 2025 | Latency-Aware Joint Task Offloading and Energy Control for Cooperative Mobile Edge ComputingabstractIn the application of the Internet of Things (IoT), existing cloud edge collaboration technologies face the problem of poor coordination of heterogeneous resources. In this article, we proposeCFEMC, which is a novelCloud-Fog-EdgeMulti-layerCollaboration resource scheduling framework for IoT. First, we design a collaborative resource scheduling framework based on semi-distributed artificial intelligence. It can achieve collaborative optimization of cloud/edge computing resource allocation under the constraints of high reliability and low latency. Second, we present a workflow applications scheduling strategy based on the proposed collaborative resource scheduling framework. This can solve the problem of unstable computing performance and transmission bandwidth during the scheduling process. Finally, the extensive and real data supported simulation results show thatCFEMChas advantages in terms of energy consumption, delay and throughput compared with other benchmark strategies. Against CEC Hu et al. 2023 and PSO Zeng et al. 2022, the average throughput increases by 16.37% and 24.21%, and the total queuing delay decreases by 54.23% and 58.12%, respectively. Weibei Fan, Fu Xiao 0001, Xiaobai Chen, Shui Yu 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | Audio-Visual Aware Foveated RenderingabstractWith the increasing complexity of geometry and rendering effects in virtual reality (VR) scenes, existing foveated rendering methods for VR head-mounted displays (HMDs) struggle to meet users' demands for VR scene rendering with high frame rates ($\geq 60fps$≥60fps for rendering binocular foveated images in VR scenes containing over 50 m triangles). Current research validates that auditory content affects the perception of the human visual system (HVS). However, existing foveated rendering methods primarily model the HVS's eccentricity-dependent visual perception ability on the visual content in VR while ignoring the impact of auditory content on the HVS's visual perception. In this article, we introduce an auditory-content-based perceived rendering quality analysis to quantify the impact of visual perception under different auditory conditions in foveated rendering. Based on the analysis results, we propose an audio-visual aware foveated rendering method (AvFR). AvFR first constructs an audio-visual feature-driven perception model that predicts the eccentricity-based visual perception in real time by combining the scene's audio-visual content, and then proposes a foveated rendering cost optimization algorithm to adaptively control the shading rate of different regions with the guidance of the perception model. In complex scenes with visual and auditory content containing over 1.17 m triangles, AvFR renders high-quality binocular foveated images at an average frame rate of 116$fps$fps. The results of the main user study and performance evaluation validate that AvFR achieves significant performance improvement (up to 1.4× speedup) without lowering the perceived visual quality compared with the state-of-the-art VR-HMD foveated rendering method. Xuehuai Shi, Jian Wu 0033, Jieming Yin, Xiaobai Chen, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | CompoundEye: A 0.24-4.17 TOPS Scalable Multi-Node DNN Processor for Image RecognitionabstractThis paper proposes a scalable DNN processor that can be flexibly reconfigured to maximize inference efficiency on a wide range of DNN models. The processor consists of 18 computing nodes with various precision modes support. To improve the computation throughput, we propose a sub-image parallelization strategy, where the original input image is divided into multiple sub-images and computed on multiple nodes in parallel. In addition, the cross-layer pipeline is implemented to improve resource utilization. The proposed processor is implemented in 28nm CMOS technology and achieves a peak performance of 4.17 TOPS and an energy efficiency of 2.08 TOPS/W. Xiaobai Chen, Qiurun Hu, Fu Xiao 0001, Jieming Yin |
ISCAS | 1 |
| 2023 | Disjoint Paths Construction and Fault-Tolerant Routing in BCube of Data Center NetworksabstractBCube is a promising structure of data center network, as it can significantly improve the performance of typical applications. With the expansion of network scale and increasement of complexity, reliability and stability of networks have become more essential. In this paper, we study the fault-tolerant routings in BCube. First, we design a fault-tolerant routing algorithm based on node disjoint multi-paths. The proposed multi-path routing has stronger fault tolerance, since each path has no other common nodes except the source node and the destination node. Second, we investigate an effective fault-tolerant routing based on routing capabilities algorithm for BCube. The proposed algorithm has higher fault tolerance and success rate of finding feasible routes, since it does not limit the faults number. Third, we present an adaptive path finding algorithm for establishing virtual links between any two nodes in BCube, which can shorten the diameter of BCube. Extensive simulation results show that the proposed routing scheme outperforms the existing popular algorithms. Compared with the state-of-the-art fault-tolerant routing algorithms, the proposed algorithm has a 21.5% to 25.3% improvement on both throughput and packet arrival rate. Meanwhile, it reduces the average latency of 18.6% and the maximum latency of 23.7% in networks. Weibei Fan, Fu Xiao 0001, Xiaobai Chen, Shui Yu 0001 |
IEEE Trans. Computers | 4 |
| 2021 | A 2.44 Tops/W Heterogeneous DCNN Inference/Training Processor for Embedded SystemabstractSince Deep Convolutional Neural Network (DCNN) training involves complex computations and data transmissions, the previous DCNN processors hard to achieve ideal energy efficiency. This paper proposed a DCNN processor supports both inference and training for the embedded system. The processor contains three heterogeneous cores to provide distinct computation patterns and dataflow for different training phases. In addition, since inference takes up more than 90% of the workload of the DCNN application, the three cores of the processor can be reconfigured to efficiently support the inference to achieve leading resources Utilization. The processor is fabricated in 55nm CMOS technology, post-layout simulation shows the processor achieving 1.36 Tops/w energy efficiency for training and 2.44 Tops/w for the inference. Xiaobai Chen, Weibei Fan, Yong Xie 0003, Fu Xiao 0001 |
ISCAS | 1 |
| 2021 | Cybersecurity protection on in-vehicle networks for distributed automotive cyber-physical systems: State-of-the-art and future challengesabstractAbstract The ever‐evolving trip mode of human being leads the automobiles moving toward connected, autonomous, sharing, and electrified vehicles rapidly. But the connection introduces new cybersecurity problems on in‐vehicle networks, which poses great challenges for safety guarantee of distributed automotive cyber‐physical systems. This article first analyzes the cybersecurity vulnerabilities and defines the security requirements for in‐vehicle networks, and then introduces the architecture evolution of in‐vehicle network. Based on the definition on architecture of in‐vehicle networks, this article defines a security protection framework for it. And then, it surveys the state‐of‐the‐art works for availability protection, integrity protection, and confidentiality protection of in‐vehicle networks, respectively, and detailed analysis and comparisons are given about the proposed cybersecurity protection mechanisms. Finally, it summarizes the future challenges for cybersecurity protection of in‐vehicle networks, and proposes possible solutions for these challenges. Yong Xie 0003, Jian Zhou 0009, Xiaobai Chen, Fu Xiao 0001 |
Softw. Pract. Exp. | 5 |
| 2021 | An MTJ-Based Asynchronous System With Extremely Fine-Grained Voltage ScalingabstractIn this work, we present an asynchronous MTJ-CMOS hybrid system with extremely fine-grained voltage scaling (EFGVS) technique. The supply voltage of the system is turned on/off by asynchronous bundled-data handshake signals. The MTJ write circuit is variation robust, self-terminated and redundant-write preventing. Besides, EFGVS and asynchronous data driven handshake enhance the timing robustness by removing the matching delay elements and reducing timing assumptions. The completion detection circuitry is also simplified. A RISC processor with proposed techniques is designed and fabricated by 55nm technology. The voltage supplies are turned on/off within tens of picoseconds. The energy for writing MTJs is 0.45pJ. The tolerance of the minimum TMR due to process variation is 75%. Sleep mode leakage power can be reduced over 10× by powering off modules with the Break-even Time of 23.6 ns. Because of the extremely fine-grained voltage scaling, no more than 20% of the modules are powered on during the execution. Ningyuan Yin, Baofa Huang, Xiaobai Chen, Zhiyi Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Efficient Virtual Network Embedding of Cloud-Based Data Center Networks into Optical NetworksabstractThe demand for data center bandwidth has exploded due to the continuous development of cloud computing, causing the use of network resources close to saturation. Optical network has become an encouraging technology for many burgeoning networks and parallel/distributed computing applications because of its huge bandwidth. This article focuses on efficient embedding of data centers into optical networks, which aims to reduce complexity of the network topology by using the parallel transmission characteristics of optical fiber. We first present a novel virtual network embedding (VNE) mathematical model used for optical data center networks. Then we derive a priority of location VNE algorithm according to node proximity sensing and path comprehensive evaluation. Furthermore, we propose routing and wavelength assignment for DCNs into optical networks, and identify the lower bound of the required number of wavelengths. Extensive evaluations show that the proposed embedding algorithm can reduce the average waiting time of virtual network requests by 20 percent, increase the request acceptance rate and revenue-overhead ratio by 13 percent, as compared to the latest VNE algorithm. Weibei Fan, Fu Xiao 0001, Xiaobai Chen, Lei Cui 0006, Shui Yu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | A 68-mw 2.2 Tops/w Low Bit Width and Multiplierless DCNN Object Detection Processor for Visually Impaired PeopleabstractDeep convolutional neural network (DCNN) object detection is a powerful solution in visual perception, but it requires huge computation and communication costs. We proposed a fast and low-power always-on object detection processor that allows visually impaired people to understand their surroundings. We designed an automatic DCNN quantization algorithm that successfully quantizes the data to 8-bit fix-points with 32 values and uses 5-bit indexes to represent them, reducing hardware cost by over 68% compared to the 16-bit DCNN, with negligible accuracy loss. A specific hardware accelerator is designed, which uses reconfigurable process engines to realize multi-layer pipelines to significantly reduce or eliminate the off-chip temporary data transfer. A lookup table is used to implement all multiplications in convolutions to reduce the power significantly. The design is fabricated in SMIC 55-nm technology, and the post-layout simulation shows only 68-mw power at 1.1-v voltage with 155 Go/s performance, achieving 2.2 Top/w energy efficiency. Xiaobai Chen, Jinglong Xu, Zhiyi Yu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | A Flexible and Energy-Efficient Convolutional Neural Network Acceleration With Dedicated ISA and Accelerator
Xiaobai Chen, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Schelling points on 3D surface meshesabstractThis paper investigates "Schelling points" on 3D meshes, feature points selected by people in a pure coordination game due to their salience. To collect data for this investigation, we designed an online experiment that asked people to select points on 3D surfaces that they expect will be selected by other people. We then analyzed properties of the selected points, finding that: 1) Schelling point sets are usually highly symmetric, and 2) local curvature properties (e.g., Gauss curvature) are most helpful for identifying obvious Schelling points (tips of protrusions), but 3) global properties (e.g., segment centeredness, proximity to a symmetry axis, etc.) are required to explain more subtle features. Based on these observations, we use regression analysis to combine multiple properties into an analytical model that predicts where Schelling points are likely to be on new meshes. We find that this model benefits from a variety of surface properties, particularly when training data comes from examples in the same object class. Xiaobai Chen, Abulhair Saparov, Bill Pang, Thomas A. Funkhouser |
ACM Trans. Graph. | 1 |
| 2010 | Möbius Transformations For Global Intrinsic Symmetry AnalysisabstractAbstract The goal of our work is to develop an algorithm for automatic and robust detection of global intrinsic symmetries in 3D surface meshes. Our approach is based on two core observations. First, symmetry invariant point sets can be detected robustly using critical points of the Average Geodesic Distance (AGD) function. Second, intrinsic symmetries are self‐isometries of surfaces and as such are contained in the low dimensional group of Möbius transformations. Based on these observations, we propose an algorithm that: 1) generates a set of symmetric points by detecting critical points of the AGD function, 2) enumerates small subsets of those feature points to generate candidate Möbius transformations, and 3) selects among those candidate Möbius transformations the one(s) that best map the surface onto itself. The main advantages of this algorithm stem from the stability of the AGD in predicting potential symmetric point features and the low dimensionality of the Möbius group for enumerating potential self‐mappings. During experiments with a benchmark set of meshes augmented with human‐specified symmetric correspondences, we find that the algorithm is able to find intrinsic symmetries for a wide variety of object types with moderate deviations from perfect symmetry. Vladimir G. Kim, Yaron Lipman, Xiaobai Chen, Thomas A. Funkhouser |
Comput. Graph. Forum | 3 |
| 2010 | Fuzzy Geodesics and Consistent Sparse Correspondences For: eformable ShapesabstractAbstract A geodesic is a parameterized curve on a Riemannian manifold governed by a second order partial differential equation. Geodesics are notoriously unstable: small perturbations of the underlying manifold may lead to dramatic changes of the course of a geodesic. Such instability makes it difficult to use geodesics in many applications, in particular in the world of discrete geometry. In this paper, we consider a geodesic as the indicator function of the set of the points on the geodesic. From this perspective, we present a new concept called fuzzy geodesics and show that fuzzy geodesics are stable with respect to the Gromov‐Hausdorff distance. Based on fuzzy geodesics, we propose a new object called the intersection configuration for a set of points on a shape and demonstrate its effectiveness in the application of finding consistent correspondences between sparse sets of points on shapes differing by extreme deformations. Jian Sun 0002, Xiaobai Chen, Thomas A. Funkhouser |
Comput. Graph. Forum | 2 |
| 2010 | Symmetry factored embedding and distanceabstractWe introduce the Symmetry Factored Embedding (SFE) and the Symmetry Factored Distance (SFD) as new tools to analyze and represent symmetries in a point set. The SFE provides new coordinates in which symmetry is "factored out," and the SFD is the Euclidean distance in that space. These constructions characterize the space of symmetric correspondences between points -- i.e., orbits. A key observation is that a set of points in the same orbit appears as a clique in a correspondence graph induced by pairwise similarities. As a result, the problem of finding approximate and partial symmetries in a point set reduces to the problem of measuring connectedness in the correspondence graph, a well-studied problem for which spectral methods provide a robust solution. We provide methods for computing the SFE and SFD for extrinsic global symmetries and then extend them to consider partial extrinsic and intrinsic cases. During experiments with difficult examples, we find that the proposed methods can characterize symmetries in inputs with noise, missing data, non-rigid deformations, and complex symmetries, without a priori knowledge of the symmetry group. As such, we believe that it provides a useful tool for automatic shape analysis in applications such as segmentation and stationary point detection. Yaron Lipman, Xiaobai Chen, Ingrid Daubechies, Thomas A. Funkhouser |
ACM Trans. Graph. | 2 |
| 2009 | A benchmark for 3D mesh segmentationabstractThis paper describes a benchmark for evaluation of 3D mesh segmentation salgorithms. The benchmark comprises a data set with 4,300 manually generated segmentations for 380 surface meshes of 19 different object categories, and it includes software for analyzing 11 geometric properties of segmentations and producing 4 quantitative metrics for comparison of segmentations. The paper investigates the design decisions made in building the benchmark, analyzes properties of human-generated and computer-generated segmentations, and provides quantitative comparisons of 7 recently published mesh segmentation algorithms. Our results suggest that people are remarkably consistent in the way that they segment most 3D surface meshes, that no one automatic segmentation algorithm is better than the others for all types of objects, and that algorithms based on non-local shape features seem to produce segmentations that most closely resemble ones made by humans. Xiaobai Chen, Aleksey Golovinskiy, Thomas A. Funkhouser |
ACM Trans. Graph. | 1 |