EDBT 2026 Demo / reviewers in the wild / expert
Jie Lou
dblp:32/5076
· DBLP profile ↗
20ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConfASR: A Conformer Block Accelerator for Speech Recognition Optimized for Edge DevicesabstractAttention-based neural networks, like transformers, have significantly improved automatic speech recognition (ASR). Adding convolution operations to transformers results in conformers, which enable better learning of local dependencies and reduce word error rate. We introduce ConfASR, the first conformer block accelerator designed for efficient ASR inference on edge devices. Our system is optimized both in terms of algorithms and hardware to support all transformer operations and additional features required for conformers, including depthwise-separable convolution and learned positional encoding. We propose a hardware-friendly normalization, shared scaling factors for non-linear functions, and an efficient dataflow with a shared MAC array that keeps all activations on chip. Implemented in a 22 nm FDSOI technology, ConfASR operates at 250 MHz with a power consumption of 359 mW, and a die area of $1.19 \mathrm{~mm}^{2}$. It performs over 900 times faster than necessary for real-time streaming requirements. This makes the architecture suitable not only for ASR but also for other transformer-based applications. ConfASR reduces latency by over $4 \times$ and power consumption by $16 \times$ during real-time use compared to previous solutions, while supporting more functionality. Malte Wabnitz, Max Nilovic, Finn Scholz, Dominik Friedrich, Christian Lanius, Jie Lou, Tobias Gemmeke |
ASP-DAC | 6 |
| 2026 | GUPrecision: Group-Wise Uniform Precision Accelerator for Depthwise Separable Convolution using Hardware-Algorithm Co-DesignabstractQuantization and pruning are effective techniques for reducing neural network size and improving energy efficiency. Although fixed word length networks are well-suited for hardware acceleration, mixed-precision and pruned networks still suffer from efficient hardware support. Depthwise separable convolution (DSC) has become a key building block for resource-constrained devices; however, applying quantization and pruning to DSC models remains challenging. To address these challenges, we propose GUPrecision, a group-wise mixed-precision uniform quantization framework that inherently supports network pruning while enabling efficient hardware realization. GUPrecision achieves hardware compatibility by dividing channels into subgroups with a fixed total bit budget. Within each group, the mixed-precision multipliers in the PE array can dynamically adapt to varying word lengths using simple shifting and multiplexing operations. We evaluated GUPrecision on the MobileNetV1 model and implemented it using GlobalFoundries 22 nm FDSOI technology. The DSC accelerator operates at 1 GHz and 0.8 V after signoff, occupying an area of 0.71 mm2 . At 75% effective sparsity, it achieves a peak energy efficiency of 17.4 TOPS/W, with a corresponding throughput of 8136 GOPS and an area efficiency of 11459 GOPS/mm2. Jie Lou, Malte Wabnitz, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Federated learning ownership verification with fixed length watermarks using hash-based message authentication code
Suxia Zhu, Jie Lou, Guanglu Sun |
J. Inf. Secur. Appl. | 2 |
| 2025 | Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from ScratchabstractXueru Wen, Jie Lou, Zichao Li, Yaojie Lu, XingYu XingYu, Yuqiu Ji, Guohai Xu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Debing Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xueru Wen, Jie Lou, Yaojie Lu 0001, XingYu, Yuqiu Ji, Guohai Xu, Ben He 0001, Xianpei Han, Le Sun 0001, Debing Zhang |
ACL (1) | 2 |
| 2025 | A 22nm 96.83-TOPS/W Time-Domain Compute-in-Memory Engine Utilizing Mixed-Fidelity for Edge-AI Applications
Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?abstractReward Models (RMs) are crucial for aligning language models with human preferences.
Currently, the evaluation of RMs depends on measuring accuracy against a validation set of manually annotated preference data.
Although this method is straightforward and widely adopted, the relationship between RM accuracy and downstream policy performance remains under-explored.
In this work, we conduct experiments in a synthetic setting to investigate how differences in RM measured by accuracy translate into gaps in optimized policy performance.
Our findings reveal that while there is a weak positive correlation between accuracy and downstream performance, policies optimized towards RMs with similar accuracy can exhibit quite different performance.
Moreover, we discover that the way of measuring accuracy significantly impacts its ability to predict the final policy performance.
Through the lens of the Regressional Goodhart effect, we recognize that accuracy, when used for measuring RM quality, can fail to fully capture the potential RM overoptimization.
This underscores the inadequacy of relying solely on accuracy to reflect their impact on policy optimization. Xueru Wen, Jie Lou, Yaojie Lu 0001, XingYu, Ben He 0001, Xianpei Han, Debing Zhang, Le Sun 0001 |
ICLR | 2 |
| 2025 | The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward ModelsabstractMultimodal Reward Models (MM-RMs) are crucial for aligning Large Language Models (LLMs) with human preferences, particularly as LLMs increasingly interact with multimodal data. However, we find that MM-RMs trained on existing datasets often struggle to generalize to out-of-distribution data due to their reliance on unimodal spurious correlations, primarily text-only shortcuts within the training distribution, which prevents them from leveraging true multimodal reward functions. To address this, we introduce a Shortcut-aware MM-RM learning algorithm that mitigates this issue by dynamically reweighting training samples, shifting the distribution toward better multimodal understanding, and reducing dependence on unimodal spurious correlations. Our experiments demonstrate significant improvements in generalization, downstream task performance, and scalability, establishing a more robust framework for multimodal reward modeling. Our source code is provided on https://github.com/alignrm/Generalizable-MM-RM. Xueru Wen, Jie Lou, Yuqiu Ji, Yaojie Lu 0001, Xianpei Han, Debing Zhang, Le Sun 0001 |
ICML | 3 |
| 2025 | An All-Digital Time-Domain Compute-in-Memory Engine for Convolutional Neural Networks in 22nmabstractThis paper presents a standard cell (SC) based time-domain compute-in-memory (TDCIM) macro for convolutional neural networks (CNNs), supporting 4b×4b multiplication. A basic cell is proposed to enable bitwise multiplication with 1-bit weights and 2-bit activations, along with double-edge computing. An 8-stage successive approximation register time-to-digital converter (SAR-TDC) is employed to convert time-domain signals into the digital domain. We leverage the inherent features of the network to enhance throughput and have fabricated the TDCIM macro in a 22nm technology. We present the measured delay and variation of the basic cell, the integral nonlinearity (INL) of the TDC, and the total computation error. The proposed macro achieves an energy efficiency of 98.44 TOPS/W at 0.55V for 4b-input and 4b-weight MAC computations. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ISCAS | 1 |
| 2024 | Semantic-focused Patch Tokenizer with Multi-branch Mixer for Visual Place RecognitionabstractVisual Place Recognition (VPR) is critical for navigation and loop closure in autonomous driving tasks, mitigating the impact of shift errors caused by dynamic changes in the environment. Due to the limited ability of backbone networks and extreme environmental changes, current methods fail to capture foundational semantic details that include the distinctive attributes for unique place identification. To address this problem, we propose a new visual token-guided VPR framework that contains a semantic-focused patch tokenizer and a multi-branch Mixer. To mitigate the inference from place-unrelated objects, the semantic-focused patch tokenizer exploits attention-based channel selection and spatial partition, which efficiently captures important semantic information within the channels and preserve spatial relationships among the backbone features. To extract abstract features with spatial structure information, the multi-branch Mixer utilizes a multi-branch structure to aggregate local and global position information, improving the robustness of global representations to environmental changes. Experimental results demonstrate that our method outperforms state-of-the-art methods, achieving 85.3% Recall@1 on the MSLS val dataset and 59.1% Recall@1 on the Nordland dataset when using ResNet18 as the backbone. Zhenyu Xu 0014, Ziliang Ren, Qieshi Zhang, Jie Lou, Dacheng Tao, Jun Cheng 0002 |
ICRA | 4 |
| 2024 | An Energy Efficient All-Digital Time-Domain Compute-in-Memory Macro Optimized for Binary Neural NetworksabstractThe deployment of neural networks on edge devices has created a growing need for energy-efficient computing. In this paper, we propose an all-digital standard cell-based time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) that is compatible with commercial digital design flow. The TDCIM macro utilizes multiple computing chains that share one threshold chain, and supports double-edge operation, parallel computing and data reuse. Time-domain wave-pipelining technique is introduced to enhance throughput while preserving accuracy. Regular placement (RP) and custom routing (CR) are employed during place and route (P&R) to reduce systematic variations. We show computing delay, POOL computation accuracy, and network test accuracy at different voltages, indicating that the proposed TDCIM macro can maintain high accuracy under PVT variations. We implemented two versions of the TDCIM macro in 22nm FDSOI technology using foundry-provided delay cells DLY40 and DLY60, respectively. At a voltage of 0.5V, the TDCIM macro achieved an energy efficiency of 1.2 (1.05) POPS/W for DLY40 (DLY60), while maintaining a baseline accuracy of 98.9% on the MNIST dataset for both designs. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Universal Information Extraction as Unified Semantic MatchingabstractThe challenge of information extraction (IE) lies in the diversity of label schemas and the heterogeneity of structures. Traditional methods require task-specific model design and rely heavily on expensive supervision, making them difficult to generalize to new schemas. In this paper, we decouple IE into two basic abilities, structuring and conceptualizing, which are shared by different tasks and schemas. Based on this paradigm, we propose to universally model various IE tasks with Unified Semantic Matching (USM) framework, which introduces three unified token linking operations to model the abilities of structuring and conceptualizing. In this way, USM can jointly encode schema and input text, uniformly extract substructures in parallel, and controllably decode target structures on demand. Empirical evaluation on 4 IE tasks shows that the proposed method achieves state-of-the-art performance under the supervised experiments and shows strong generalization ability in zero/few-shot transfer settings. Jie Lou, Yaojie Lu 0001, Dai Dai, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
AAAI | 1 |
| 2023 | Learning In-context Learning for Named Entity RecognitionabstractJiawei Chen, Yaojie Lu, Hongyu Lin, Jie Lou, Wei Jia, Dai Dai, Hua Wu, Boxi Cao, Xianpei Han, Le Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiawei Chen 0011, Yaojie Lu 0001, Jie Lou, Dai Dai, Hua Wu 0003, Boxi Cao, Xianpei Han, Le Sun 0001 |
ACL (1) | 4 |
| 2023 | Scalable Time-Domain Compute-in-Memory BNN Engine with 2.06 POPS/W Energy Efficiency for Edge-AI DevicesabstractTime-domain (TD) computing has attracted attention for its high computing efficiency and suitability for applications on energy-constrained edge devices. In this paper, we present a time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) realized by standard as well as custom delay cells. Multiply-and-accumulate (MAC) operations, batch normalization (BN) and binarization (Bin) are all processed in the time-domain, avoiding costly digital domain post-processing. In addition, it supports flexible mapping for different kernel sizes, achieving 100% utilization. Starting from a standard cell-based implementation, we propose two custom cells that provide interesting trade-offs between energy efficiency, area and accuracy. The two proposed custom designs can achieve 1.5 and 2.06 POPS/W energy efficiencies at 0.5V and 0.6V with less cell area while maintaining model test accuracy. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | Automatic Generation of Structured Macros Using Standard Cells ‒ Application to CIMabstractRegularity can be exploited to efficiently describe, place and route logic blocks with a repetitive structure. We present a design flow to automatically generate regular, standard-cell based designs, which can be seamlessly integrated into a traditional digital flow in commercial EDA software. The generated arrays can be seamlessly integrated into a “sea of gates”. with no guard-rings or keep-out areas. The flow takes a description of a regular design as an input and generates netlist, placement, constraints, routing, initial parasitics estimates and timing information. We show that, in example designs, the run-time of EDA tooling is up to 2.5x faster, reduces the critical path by 47%, reduces the metal utilization by 45% and achieves a utilization of 93%. Christian Lanius, Jie Lou, Johnson Loh, Tobias Gemmeke |
ISLPED | 2 |
| 2023 | RHINE: Robust and High-performance Internet Naming with E2E Authenticity
Huayi Duan, Rubén Fischer, Jie Lou, Si Liu 0003, David A. Basin, Adrian Perrig |
NSDI | 3 |
| 2023 | An Energy-Efficient and Area-Efficient Depthwise Separable Convolution Accelerator with Minimal On-Chip Memory AccessabstractDepthwise separable convolution (DSC) has emerged as a crucial building block for developing lightweight convolutional neural networks (CNNs). In this paper, we present a hardware accelerator for DSC that enables 100% utilization of the processing element (PE) array for depthwise convolution (DWC) and achieves up to 98% utilization for pointwise convolution (PWC), while also reducing latency. By partitioning the input feature map (ifmap) SRAM of the DWC into three banks, we minimize memory access and maximize data reuse. The input activations and weights only need to be loaded once from SRAM to PE for both DWC and PWC. Additionally, to support efficient operations across different layers, we present a layerwise matching method. The proposed DSC accelerator is implemented in 22nm FDSOI technology and validated using MobileNetV1 on the CIFAR10 dataset. The post-layout results demonstrate that the proposed accelerator can operate at 1GHz and achieve an energy efficiency of 5.07 (3.96) TOPS/W and an area efficiency of 519.2 (461.52) GOPS/mm2for DWC (PWC) at 0.8V. After scaling the supply voltage down to 0.5V, the energy efficiency for the proposed accelerator increases to 13.64 TOPS/W for DWC and 10.64 TOPS/W for PWC, respectively. Jie Lou, Christian Lanius, Florian Freye, Johnson Loh, Tobias Gemmeke |
VLSI-SoC | 2 |
| 2023 | ArkGPU: enabling applications' high-goodput co-location execution on multitasking GPUs
Jie Lou, Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Ninghui Sun |
CCF Trans. High Perform. Comput. | 1 |
| 2022 | Understanding the deep structure use of mobile phones - an attachment perspectiveabstractUsers rely heavily on mobile phones as multi-purpose devices to support all kinds of tasks. This study conceptualises mobile phone deep structure use and aims at unveiling such behaviour. Towards this end, this study goes beyond the traditional technology usage theories based on rationality and instrumentality. It is replaced by a relationship-based theory, the attachment theory, to investigate how users’ attachment to mobile phones is formed and how such attachment influences mobile phone deep structure use. A survey approach was adopted to test the research model empirically. The results showed that mobile functional dependence and mobile identity constituted the two dimensions of mobile attachment, which significantly affected mobile phone deep structure use. Mobile functional dependence was formed based on the two enabling factors (perceived functional flexibility and perceived mobility). Mobile identity was formed based on the two enriching factors (self-expressive symbolism and categorical symbolism). Both mobile functional dependence and mobile identity were influenced by the gratifying factor, design aesthetics. This study has made contribution to the literature of mobile phone usage by clarifying the special characteristics of mobile phone usage context, proposing a new theoretical perspective and constructing a nomological network of the mobile attachment construct. Jie Lou, Lingzhao Deng |
Behav. Inf. Technol. | 1 |
| 2015 | Learning to suggest questions in social media
Tom Chao Zhou, Michael R. Lyu, Irwin King, Jie Lou |
Knowl. Inf. Syst. | 4 |
| 2013 | Contributing high quantity and quality knowledge to online Q&A communitiesabstractThis study investigates the motivational factors affecting the quantity and quality of voluntary knowledge contribution in online Q&A communities. Although previous studies focus on knowledge contribution quantity, this study regards quantity and quality as two important, yet distinct, aspects of knowledge contribution. Drawing on self‐determination theory, this study proposes that five motivational factors, categorized along the extrinsic‐intrinsic spectrum of motivation, have differential effects on knowledge contribution quantity versus quality in the context of online Q&A communities. An online survey with 367 participants was conducted in a leading online Q&A community to test the research model. Results show that rewards in the reputation system, learning, knowledge self‐efficacy, and enjoy helping stand out as important motivations. Furthermore, rewards in the reputation system, as a manifestation of the external regulation, is more effective in facilitating the knowledge contribution quantity than quality. Knowledge self‐efficacy, as a manifestation of intrinsic motivation, is more strongly related to knowledge contribution quality, whereas the other intrinsic motivation, enjoy helping, is more strongly associated with knowledge contribution quantity. Both theoretical and practical implications are discussed. Jie Lou, Yulin Fang, Kai H. Lim, Jerry Zeyu Peng |
J. Assoc. Inf. Sci. Technol. | 1 |