Yiou Wang

dblp:35/8155 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 1 since 2021Systems, architecture and hardware · 7 · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 PMM3D: a transformer-based monocular 3D detector with parallel multi-time inquiry and mixup enhancement
Tongzhou Zhang 0001, Wei Zhou 0051, Yiou Wang, Gang Wang 0013
Expert Syst. Appl.4
2026 Efficient SRAM-PIM Co-Design by Joint Exploration of Value-Level and Bit-Level Sparsity
abstract
Processing-in-memory (PIM) architectures mitigate the Von Neumann bottleneck by integrating computation units into memory arrays. Among PIM architectures, digital SRAMPIM has become a prominent approach, directly integrating digital logic within the SRAM array. However, the rigid crossbar architecture and full array activation pose challenges in efficiently utilizing value-level sparsity. Moreover, neural network models exhibit a high proportion of zero bits within non-zero values, which remain underutilized due to architectural constraints. To overcome these limitations, we present Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework to harness both value-level and bit-level sparsity. At the algorithm level, our hybrid-grained pruning technique, combined with a novel sparsity pattern, enables effective sparsity management. Architecturally, DB-PIM incorporates a sparse network and customized digital SRAM-PIM macros, including input pre-processing unit (IPU), dyadic block multiply units (DBMUs), and Canonical Signed Digit (CSD)-based adder trees. It circumvents structured zero values in weights and bypasses unstructured zero bits within non-zero weights and block-wise all-zero bit columns in input features. As a result, the DBPIM framework skips a majority of unnecessary computations, thereby driving significant gains in computational efficiency. Experimental results demonstrate that our DB-PIM framework achieves up to 8.01× speedup and 85.28% energy savings, significantly boosting computational efficiency in digital SRAMPIM systems.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM Architectures
abstract
Digital Compute-in-Memory (CIM) architectures have shown great promise in Deep Neural Network (DNN) acceleration by effectively addressing the “memory wall” bottleneck. However, the development and optimization of digital CIM accelerators are hindered by the lack of comprehensive tools that encompass both software and hardware design spaces. Moreover, existing design and evaluation frameworks often lack support for the capacity constraints inherent in digital CIM architectures. In this paper, we present CIMFlow, an integrated framework that provides an out-of-the-box workflow for implementing and evaluating DNN workloads on digital CIM architectures. CIMFlow bridges the compilation and simulation infrastructures with a flexible instruction set architecture (ISA) design, and addresses the constraints of digital CIM through advanced partitioning and parallelism strategies in the compilation flow. Our evaluation demonstrates that CIMFlow enables systematic prototyping and optimization of digital CIM architectures across diverse configurations, providing researchers and designers with an accessible platform for extensive design space exploration.
Yingjie Qi, Jianlei Yang 0001, Yiou Wang, Dayu Wang, Cenlin Duan, Xiaolin He, Weisheng Zhao 0001
DAC3
2025 Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
abstract
Graph Neural Networks (GNNs) have been widely adopted due to their strong performance. However, GNN training often relies on expensive, high-performance computing platforms, limiting accessibility for many tasks. Profiling of representative GNN workloads indicates that substantial efficiency gains are possible on resource-constrained devices by fully exploiting available resources. This paper introduces$\mathrm{A}^{3} \text{GNN}$, a framework for Affordable, Adaptive, and Automatic GNN training on heterogeneous CPU-GPU platforms. It improves resource usage through locality-aware sampling and fine-grained parallelism scheduling. Moreover, it leverages reinforcement learning to explore the design space and achieve pareto-optimal trade-offs among throughput, memory footprint, and accuracy. Experiments show that$\mathrm{A}^{3}$GNN can bridge the performance gap, allowing seven Nvidia 2080Ti GPUs to outperform two A100 GPUs by up to$1.8 \times$in throughput with minimal accuracy loss.
Yingjie Qi, Yiou Wang, Han Wan, Jianlei Yang 0001, Chunming Hu
ICCD4
2025 Mixanthropy: Holographic Metamorphic Clouds
Meichun Cai, Yiou Wang
ACM Multimedia2
2024 Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
abstract
Bit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computational efficiency. Yet, traditional digital SRAM-PIM architecture, limited by rigid crossbar architecture, struggles to effectively exploit this unstructured sparsity. To address this challenge, we propose Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework. First, we propose an algorithm coupled with a distinctive sparsity pattern, termed a dyadic block (DB), that preserves the random distribution of non-zero bits to maintain accuracy while restricting the number of these bits in each weight to improve regularity. Architecturally, we develop a custom PIM macro that includes dyadic block multiplication units (DBMUs) and Canonical Signed Digit (CSD)-based adder trees, specifically tailored for Multiply-Accumulate (MAC) operations. An input pre-processing unit (IPU) further refines performance and efficiency by capitalizing on block-wise input sparsity. Results show that our proposed co-design framework achieves a remarkable speedup of up to 7.69× and energy savings of 83.43%.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
DAC3
2024 DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-Memory
abstract
Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively.
Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 Moving Architecture, Animated Maze: The Intertextuality Between Player and Environment Agents
abstract
The maze, a classic architectural type of post-functional nature, is reinvented through the contemporary lens of video game. With novel analytical insights and computing methodologies on the system-unit relationship of a maze, we design and develop a Moving Maze that moves its parts methodically in response to the player’s movement. A disorienting and adaptive system composed of identical parts, the Moving Maze is deconstructed into the non-subdivisible unit, which can propagate into a field through replication and orthogonal rotation. The game generates unit-to-system interactive outcomes with fragmental movements using gamer-relational rules. In achieving difficulty progression and game balance through Reinforcement Learning, the maze arouses problem-solving curiosity and immerses the player in a risk-reward structure. Centering the game mechanics on interactivity and adaptability, we enhance player engagement in this cognitive puzzle game through balanced player and environment agency. An artwork synthesizing procedural computation with gaming architecture, the Moving Maze pushes the imaginative boundary of what a maze can be and embodies the philosophy that systemic complexities arise from the simplest elements.
Yiou Wang
Creativity & Cognition1
2021 A Bibliometric Analysis of Edge Computing for Internet of Things
abstract
In recent years, with the emergence of many Internet of Things applications such as smart homes, smart city, and connected vehicles, the amount of network edge data increases rapidly. Now, edge computing for Internet of Things has attracted the research interest of many researchers. Then, a thorough analysis of the current body of knowledge in edge computing for Internet of Things is conducive to a comprehensive understanding of the research status and future trends in this field. In this paper, a bibliometric analysis of edge computing for Internet of Things was performed using the Web of Science (WoS) Core Collection dataset. The relevant literature studies published in this field were quantitatively analyzed based on a bibliometric analysis method combined with VOSviewer software, and the development history, research hotspots, and future directions of this field were studied. The research results show that the number of literature studies published in the field of edge computing for Internet of Things is on the rise over time, especially after 2017, and the growth rate is accelerating. China and USA take the lead position in the number of literature studies published. Zhang is the most productive author, and Satyanarayanan is the most influential author. IEEE Access and IEEE Internet of Things Journal are the main journals in this field. Beijing University of Posts Telecommunications has published most literature studies. Research hotspots of edge computing for Internet of Things mainly include specific problem research such as resource management, architecture research, application research, and fusion research of this field with some other fields such as artificial intelligence and 5G.
Yiou Wang, Fuquan Zhang 0001, Laiyang Liu
Secur. Commun. Networks1
2021 Parallel optimization of the ray-tracing algorithm based on the HPM model
Yiou Wang, Yu-Gang Li, Fuquan Zhang 0001
J. Supercomput.3
2021 Correction to: Parallel optimization of the ray-tracing algorithm based on the HPM model
abstract
A correction to this paper has been published: https://doi.org/10.1007/s11227-021-03680-0
Yiou Wang, Yu-Gang Li, Fuquan Zhang 0001
J. Supercomput.3
2019 A Semi-Supervised Approach for Identification of the Sections in Charge of RFQ Documents
abstract
Identification of sections in charge of a request for quotation (RFQ), a type of business-specific document that seeks an itemized list of prices for a product or service, is usually performed manually and is very time-consuming, especially in the manufacturing industry. This study presents a simple semi-supervised classification approach for automatic section identification of RFQ documents. We conceive the identification task as text classification task for different sections and introduce novel features derived from unlabeled data to enhance the performance. We evaluate the usefulness of our approach in a series of experiments on a collection of RFQ documents in the actual business operations and obtain satisfactory results for most test collections.
Izumo Hidetaka, Yiou Wang
IEEE BigData2
2018 A Japanese Corpus for Analyzing Customer Loyalty Information
Yiou Wang, Takuji Tahara
LREC1
2017 Customer Churn Prediction Using Sentiment Analysis and Text Classification of VOC
Yiou Wang, Koji Satake, Hiroshi Masuichi
CICLing (2)1
2012 Chinese Evaluative Information Analysis
Yiou Wang, Jun'ichi Kazama, Takuya Kawada, Kentaro Torisawa
COLING1
2012 Why Question Answering using Sentiment Analysis and Word Classes
Jong-Hoon Oh, Kentaro Torisawa, Chikara Hashimoto, Takuya Kawada, Stijn De Saeger, Jun'ichi Kazama, Yiou Wang
EMNLP-CoNLL7
2012 Bitext Dependency Parsing With Auto-Generated Bilingual Treebank
abstract
This paper proposes a method to improve the accuracy of bilingual texts (bitexts) dependency parsing by using an auto-generated bilingual treebank created with the help of statistical machine translation (SMT) systems. Previous bitext parsing methods use human-annotated bilingual treebanks that are costly and troublesome to obtain. In the proposed method, we use an auto-generated bilingual treebank to train the parsing models. First, an SMT system is used to translate a monolingual treebank into the target language; then, a monolingual parser for the target language is used to parse the translated sentences. Since the auto-translated sentences and auto-parsed trees in the auto-generated bilingual treebank are far from perfect, the bilingual constraints are not sufficiently reliable. To overcome this problem, we propose a method to verify the reliability of the constraints using a large amount of target monolingual and bilingual unannotated data. Finally, we design a set of effective bilingual features for parsing models on the basis of the verified constraints. We conduct the experiments using a standard test data. The experimental results show that our bitext parser significantly outperforms monolingual parsers. Moreover, our method is still able to provide improvement when we use a larger monolingual treebank containing over 50 000 sentences. We also test the proposed method with different SMT systems and the results show that our method is very robust to the noise. In particular, the proposed method can be used in a purely monolingual setting with the help of SMT. That is, it does not need the human translation of the test set as previous methods do.
Wenliang Chen, Jun'ichi Kazama, Min Zhang 0005, Yoshimasa Tsuruoka, Yiou Wang, Kentaro Torisawa, Haizhou Li 0001
IEEE Trans. Speech Audio Process.6
2011 SMT Helps Bitext Dependency Parsing
Wenliang Chen, Jun'ichi Kazama, Min Zhang 0005, Yoshimasa Tsuruoka, Yiou Wang, Kentaro Torisawa, Haizhou Li 0001
EMNLP6
2011 Improving Chinese Word Segmentation and POS Tagging with Semi-supervised Methods Using Large Auto-Analyzed Data
Yiou Wang, Jun'ichi Kazama, Yoshimasa Tsuruoka, Wenliang Chen, Kentaro Torisawa
IJCNLP1
2010 Adapting Chinese Word Segmentation for Machine Translation Based on Short Units
Yiou Wang, Kiyotaka Uchimoto, Jun'ichi Kazama, Canasai Kruengkrai, Kentaro Torisawa
LREC1
2009 An Error-Driven Word-Character Hybrid Model for Joint Chinese Word Segmentation and POS Tagging
Canasai Kruengkrai, Kiyotaka Uchimoto, Jun'ichi Kazama, Yiou Wang, Kentaro Torisawa, Hitoshi Isahara
ACL/IJCNLP4