EDBT 2026 Demo / reviewers in the wild / expert
Xinhua Chen
dblp:24/1926
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CROPHE: Cross-Operator Dataflow Optimization for Fully Homomorphic Encryption AcceleratorsabstractFully homomorphic encryption (FHE) enables the protection of data privacy at the cost of significantly higher computational demands. To alleviate its memory-bound bottlenecks, dataflow optimizations that maximize on-chip data reuse and minimize off-chip accesses could be leveraged. In this work, we exploit the opportunities of cross-operator dataflow optimizations in FHE accelerators, and propose a hardware-software co-design called CROPHE. On the hardware level, instead of overly-specialized functional units, CROPHE provisions a homogeneous and unified architecture that allows for flexible resource allocation and operator mapping. On the software level, the scheduling framework of CROPHE takes a comprehensive and systematic approach to explore various spatial and temporal data pipelining and sharing schemes across multiple operators, resulting in more efficient dataflow than prior work. We also propose novel cross-operator dataflow optimizations for the unique operators in FHE including number theoretic transforms and homomorphic rotations. The evaluation shows CROPHE significantly outperforms state-of-the-art designs by$1.77 \times$to$4.86 \times$. Xinhua Chen, Jiangbin Dong, Hongren Zheng, Tian Tang 0001, Mingyu Gao 0001 |
HPCA | 1 |
| 2026 | EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan 0003, Cong Guo 0003, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001 |
ISCA | 6 |
| 2025 | A Unified Vector Processing Unit for Fully Homomorphic EncryptionabstractFully homomorphic encryption (FHE) algorithms enable privacy-preserving computing directly on encrypted data without leaking sensitive contents, while their excessive computational overheads could be alleviated by specialized hardware accelerators. The vector architecture has been prominently used for FHE accelerators to match the underlying polynomial data structures. While most FHE operations can be efficiently supported by vector processing units, the number theoretic transform (NTT) and automorphism operators involve complex and irregular data permutations among vector elements, and thus are handled with separate dedicated hardware units in existing FHE accelerators. In this paper, we present an efficient inter-lane network design and the corresponding dataflow control scheme, in order to realize NTT and automorphism operations among the multiple lanes of a vector unit. An arbitrarily large operator is first decomposed to fit in the fixed width of the vector unit, and the required data permutation and transposition are conducted on the specialized inter-lane network. Compared to previous designs, our solution reduces the hardware resources needed, with up to 9.4x area and 6.0x power savings for only the inter-lane network, and up to 1.2 x area and 1.1 x power savings for the whole vector unit. Jiangbin Dong, Xinhua Chen, Mingyu Gao 0001 |
DATE | 2 |
| 2025 | Sliding-Window Scheduling to Exploit Hybrid-Bonding-Based Accelerators for Fully Homomorphic Encryption
Xinhua Chen, Xinglong Yu, Yifan Zhao 0007, Honglin Kuang, Jun Han 0003 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Imaging Multimode Dispersion Curves of Rayleigh and Love Waves From Three-Component Noise Recordings in Urban EnvironmentsabstractWith the advancements of three-component seismic instruments, much valuable information about source distributions and subsurface structures can be utilized to improve passive surface-wave imaging with anthropogenic seismic noise in urban environments. Current passive surface-wave methods, however, are mainly concerned with the Rayleigh waves in the vertical (Z) component, often neglecting the useful dispersion information in the radial (R) and transverse (T) components, particularly in higher modes of surface waves. So, we introduced the common-midpoint two-station analysis (CMP-TS) to extract the multimode dispersion curves of Rayleigh and Love waves from three-component noise recordings using multicomponent ambient seismic noise cross-correlations (ZZ, RR, and TT components). Results from synthetic data sets from given models show that the CMP-TS method is able to retrieve higher-mode Rayleigh waves from the RR component and improve the multimode dispersion measurements of Rayleigh waves by the summation of the ZZ and RR spectrograms. Besides, this method can extract multimode dispersion curves of Love waves with high resolution from the TT component. We applied the CMP-TS method to process three-component field data and retrieved the dispersion curves of Rayleigh waves with the fundamental and the first higher modes, as well as Love waves with the fundamental, the first, and second higher modes. The S-wave velocity model is constructed by inverting the fundamental and higher mode data and validated through borehole S-wave velocity measurements. Jingyin Pang, Xuben Wang, Jianghai Xia, Binbin Mi, Xinhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Zero-Shot Structure-Preserving Diffusion Model for High Dynamic Range Tone MappingabstractTone mapping techniques, aiming to convert high dynamic range (HDR) images to high-quality low dynamic range (LDR) images for display, play a more crucial role in real-world vision systems with the increasing application of HDR images. However, obtaining paired HDR and high-quality LDR images is difficult, posing a challenge to deep learning based tone mapping methods. To over-come this challenge, we propose a novel zero-shot tone mapping framework that utilizes shared structure knowl-edge, allowing us to transfer a pre-trained mapping model from the LDR domain to HDR fields without paired training data. Our approach involves decomposing both the LDR and HDR images into two components: structural in-formation and tonal information. To preserve the original image's structure, we modify the reverse sampling process of a diffusion model and explicitly incorporate the struc-ture information into the intermediate results. Additionally, for improved image details, we introduce a dual-control network architecture that enables different types of conditional inputs to control different scales of the output. Experimental results demonstrate the effectiveness of our approach, surpassing previous state-of-the-art methods both qualitatively and quantitatively. Moreover, our model ex-hibits versatility and can be applied to other low-level vi-sion tasks without retraining. The code is available at https://github.com/ZSDM-HDRIZero-Shot-Diffusion-HDR. Ruoxi Zhu, Shusong Xu, Peiye Liu, Sicheng Li 0001, Yanheng Lu, Dimin Niu, Zihao Liu 0015, Zihao Meng, Zhiyong Li 0016, Xinhua Chen, Yibo Fan |
CVPR | 10 |
| 2024 | Surface Wave Inversion Using a Multi-Information Fusion Neural NetworkabstractIn noninvasive near-surface investigations, with the emergence of massive seismic datasets, surface wave inversion using deep learning (DL) can efficiently attain the shear-wave velocity (Vs) model. Existed researches on DL inversion, however, cannot handle the inversion nonuniqueness effectively. The geological constraint is only reflected in their training dataset. The input of their neural networks only contains dispersion curves, and thus the density and compressional-wave velocity (Vp) need self-learning. To decrease the nonuniqueness, we propose a multi-information fusion neural network (MFNN) in which we add the Vp, density, and sensitivity as parts of the input. To verify the effectiveness of the MFNN, we used a synthetic test and field work which both contain data from six regions to conduct multi-region simultaneous inversions. We compared the inversion results of the MFNN with the results of a convolutional neural network and the ground truth. In the two experiments, although the calculated dispersion curves based on the inverted Vs model from all the methods match well with the observed ones, only the inverted Vs from the MFNN successfully reflect the actual Vs structure and locate the high-velocity layer and low-velocity layer accurately. The constraints on the Vp, density, and sensitivity thus effectively reduce the nonuniqueness of inversions. Xinhua Chen, Jianghai Xia, Jingyin Pang, Hao Zhang 0185 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2012 | A 16-pixel parallel architecture with block-level/mode-level co-reordering approach for intra prediction in 4k×2k H.264/AVC video encoderabstractIntra prediction is the most important technology in H.264/AVC intra frame encoder. But there is extremely complicated data dependency and an immense amount of computation in intra prediction process. In order to meet the requirements of real-time coding and avoid hardware waste, this paper presents a parallel and high efficient H.264/AVC intra prediction architecture which targets high-resolution (e.g. 4k×2k) video encoding applications. In this architecture, the optimized intra 4×4 prediction engine can process sixteen pixels in parallel at a slightly higher hardware cost (compared to the previous four-pixel parallel architecture). The intra 16×16 prediction engine works in parallel with intra 4×4 prediction engine. It reuses the adder-tree of Sum of Absolute Transformed Difference (SATD) generator. Moreover, in order to reduce the data-dependency in intra 4×4 reconstruction loop, a block-level and mode-level co-reordering strategy is proposed. Therefore, the performance bottleneck of H.264/AVC intra encoding can be alleviated to a great extent. The proposed architecture supports full-mode intra prediction for H.264/AVC baseline, main and extended profiles. It takes only 163 cycles to complete the intra prediction process of one macroblock (MB). This design is synthesized with a SMIC 0.13µm CMOS cell library. The result shows that it takes 61k gates and can run at 215MHz, supporting real-time encoding of 4k×2k@40fps video sequences. Huailu Ren, Yibo Fan, Xinhua Chen, Xiaoyang Zeng |
ASP-DAC | 3 |
| 2011 | A full-mode FME VLSI architecture based on 8×8/4×4 adaptive Hadamard Transform for QFHD H.264/AVC encoderabstractAdaptive Block-size Transform (ABT) has been added to H.264/AVC standard with the Fidelity Range Extension. In this paper, we apply this ABT concept to our FME design and propose a full-mode FME architecture based on 8×8/4×4 adaptive Hadamard Transform. This technique can avoid unifying all variable block-size blocks into 4×4-size blocks and improve the encoding performance. We also exploit the linearity of Hadamard Transform in quarter-pel refinement and decrease the cycles caused by the second long search process. In architecture level, we employ two interpolating engines that can support 8-pel and 4-pel input to time-share one SATD (Sum of Absolute Hadamard Transform) Generator. These strategies can increase parallelism and reduce the cycles efficiently. Besides, this design can support full modes, which guarantees the encoding performance. Experimental results show that our design can achieve real-time processing for QFHD@30fps at the operation frequency of 320MHz with 444.6K gates hardware. Jialiang Liu, Xinhua Chen, Yibo Fan, Xiaoyang Zeng |
VLSI-SoC | 2 |
| 2011 | MUX-MCM based quantization VLSI architecture for H.264/AVC high profile encoderabstractThis paper presents a hardware-efficient and high-throughput quantization implementation for H.264/AVC high profiles encoder. The constant multiplication in quantization is accomplished by time-multiplexed multiple-constant multipliers (MUX-MCM). By rational pipeline decision, the proposed design manages to achieve a high throughput at a low area cost. Synthesized with SMIC0.18µm technology, the proposed design reaches a maximum operating frequency of 250Mhz with a throughput of 1Gpixels/sec at the hardware cost of 28.56 Kgates. Jiang Ying, Xinhua Chen, Yibo Fan, Xiaoyang Zeng |
VLSI-SoC | 2 |
| 2005 | A 9.5mW 4GHz WCDMA frequency synthesizer in 0.13µm CMOSabstractA 4GHz integer-N frequency synthesizer is realized in a 0.13μm CMOS technology. It has a 400kHz reference frequency and 40kHz loop bandwidth such that 2GHz quadrature LO signals can be generated after a divide-by-two, with channel raster of 200kHz. The measured in-band phase noise is -74dBc/Hz @4kHz offset. A self-regulated charge pump is proposed to improve matching as well as charge sharing. Reference spurs are thereby kept below -55dBc over the VCO tuning voltage from rail to rail. The requirements for UMTS transceiver have been fulfilled with an overall power consumption of 9.5mW, which is the lowest reported to date. Core area of the chip is as small as 0.2mm2 Xinhua Chen, Qiuting Huang |
ISLPED | 1 |