Yi Man

dblp:20/6165 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LTTL: A Low-Overhead and Triple-Node-Upset-Tolerant Latch Design for Aerospace Applications
abstract
As the feature size of the CMOS technology keeps scaling down, the charge sharing caused by radiation is becoming more and more prominent, and the occurrence possibility of the triple-node upset (TNU) increases significantly. In this paper, we propose a low-overhead and TNU-tolerant latch (LTTL) that leverages three parallel storage cells and an output-level error interceptive module to achieve complete TNU tolerance while minimizing design overhead. The optimized structure eliminates redundant devices and employs a high-speed D-to-Q path, significantly reducing delay-area-power product (DAPP). Even any three nodes of the latch are flipped at the same time, the output of the latch can retain the original value. Simulation results not only confirm the TNU tolerance of the proposed latch but also demonstrate that the latch can provide a 57% reduction in delay, 20% reduction in area, and 62% reduction in DAPP on average compared to state-of-the-art TNU-tolerant latches.
Zikang Ma, Zhongyu Gao, Qianhui Liu, Yi Man, Huaguo Liang, Xiaoqing Wen
ACM Great Lakes Symposium on VLSI5
2026 Efficient Instruction Fusion for ASIPs: A Low-Cost Approach to Structural Hazard Control
abstract
Moore’s law is reaching saturation, while the demand for computing capacity continues to rise. Application-Specific Instruction-Set Processors (ASIPs) are therefore an essential technology for embedded processing in areas such as communication, media, gaming, and control systems, due to their high area and energy efficiency. The most effective method for ASIPs acceleration is known as instruction fusion. Instruction fusion enhances performance but also introduces challenges related to the complexity of structural hazard control. The difficulty in this research field is achieving fine-grained structural hazard control using low-cost hardware. The state-of-the-art technologies have not been able to effectively address this challenge. Firstly, this paper systematically analyzes the rationale behind instruction fusion and explores methods to optimize its performance. Secondly, based on the design of a 5G micro-base station baseband processor, we present an instruction fusion pipeline design example. Furthermore, to address the structural hazards that arise from instruction fusion, we propose a lightweight Hardware Resource Table (HRT) to address the issue and outline its benefits. These benefits include utilizing low silicon costs to achieve performance improvements and reducing on-chip program memory. According to benchmark evaluations, area efficiency improves by 23%, accompanied by a 108% increase in energy efficiency.
Xinbing Zhou, Tianlang Liu, Tiancheng Tang, Yi Man, Wei Chen 0100, Dake Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 Sayram: A Hardware-software Co-design to Accelerate Wireless Baseband Processing
abstract
Micro base stations, with limited antennas and extensive deployment, require scaled-down hardware. Software-defined radio solutions (e.g., CPU, many-core systems, GPU) offer flexibility but incur high area and power costs, while traditional DSP lacks efficient acceleration for smaller configurations. The key challenge for micro base stations is achieving minimal area and power overhead while meeting 5G requirements. This article presents a hardware-software co-designed architecture, Sayram, which minimizes overhead for 5G physical layer processing. Sayram integrates an instruction fusion mechanism, along with the compiler for simplified programming, a Vector Indirect Addressing Memory (VIAM) to minimize memory access cycles, and an improved vector register design to accelerate small-scale matrix computation, thereby improving overall processor efficiency. Operating at 1 GHz, Sayram achieves 158 GOPS with a 1.18 mm \(^2\) area, supporting 2T2R and 4T4R Physical Uplink Shared Channel (PUSCH) processing in single-core and dual-core modes, respectively. Evaluations show that Sayram’s area efficiency is 3× and 9× higher than traditional DSP and CGRA architectures, respectively, with power efficiency improvements of 44× and 6×. Sayram’s energy and area efficiency surpass CPU solutions by orders of magnitude.
Xinbing Zhou, Shaobo Shi, Shao-Han Liu, Yunxiang Tang, Tiancheng Tang, Yi Man, Dake Liu
ACM Trans. Embed. Comput. Syst.7
2025 Pipe-DBT: enhancing dynamic binary translation simulators to support pipeline-level simulation
abstract
In response to the lack of pipeline behavior modeling in Instruction-Set Simulators (ISS) and the performance limitations of Cycle-Accurate Simulators (CAS), this paper proposes Pipe-DBT, a pipeline simulation framework based on Dynamic Binary Translation (DBT). This method achieves a balance between accuracy and efficiency through two key techniques: (1) the design of a pipeline state descriptor called Pipsdep, which abstracts data hazards and resource contentions in the form of formal rules about resource occupancy and read/write behaviors, thereby avoiding low-level hardware details; (2) the introduction of a coroutine-based instruction execution flow partitioning mechanism that employs dynamic suspension/resumption to realize cycle-accurate scheduling in multi-stage pipelines. Implemented on QEMU, Pipe-DBT supports variable-length pipelines, a Very Long Instruction Word (VLIW) architecture with four-issue capability, and pipeline forwarding. Under typical DSP workloads, it achieves a simulation speed of 400–1100 KIPS, representing a 2.3 $$\times$$ improvement over Gem5 in cycle-accurate mode. Experimental results show that only modular extensions to the host DBT framework are required to accommodate heterogeneous pipeline microarchitectures, thereby providing a high-throughput simulation infrastructure for processor design. To the best of our knowledge, this is the first pipeline-level simulation model implemented on a DBT simulator.
Tiancheng Tang, Yi Man, Xinbing Zhou, Duqing Wang
Autom. Softw. Eng.2
2025 Layer-Aware Containerized Microservice Scheduling via Multiobjective Proximal Policy Optimization in Edge-Computing-Enabled IoT
abstract
Containerized microservice (MS) architecture has emerged as the preferred solution for increasingly complex IoT applications in edge computing (EC). However, containerized MS’s runtime requires frequent pulling the layer-based container image and data transfer between dependent nodes, which may incur huge traffic and latency overhead, especially in resource-limited EC-enabled IoT. Additionally, focusing only on reducing such overhead potentially harms load balance, thus degrading service performance. To address these issues, we propose an innovative layer-aware concurrent containerized MS scheduling framework to jointly trade off both consumers’ QoS (latency) and service provider’s profit (load balance). Specifically, we first model the scheduling problem as a multi-objective Markov Decision Process, fully accounting for the latency of each phase in MS’s lifecycle and various perceptible affinities related to the IoT application’s Directed Acyclic Graph. Secondly, a multi-objective DRL algorithm (MO-PPO) is proposed by reconfiguring the Actor-Critic network and devising an empirical trajectory sampling balance mechanism, enabling only a single trained model to approximate the Pareto front. Finally, extensive experiments based on real-world data traces show MO-PPO reduces the service latency by 31.8% and load imbalance by 38.7% at the optimal trade-off point compared to the baselines.
Shijun Ma, Junjie Teng, Yi Man, Yinglei Teng
IEEE Internet Things J.3
2025 Process Monitoring and Fault Prediction of Papermaking by Learning From Imperfect Data
abstract
Fault prediction is increasingly concerned in the industry due to complexity grows in the production process. Paper break, the most common process fault of papermaking, risks paper mills enormously on cost and efficiency. Data derived from papermaking workshop always involves imperfect issues of mismatching, losing, human intervention etc. which mask the inherent hints about paper break, prevent early warn. This study proposed pretreatment processes on data upon papermaking knowledge and analysis of paper breaks, exploited random forest to extract interrelated features, and establish a prediction model of paper break based on Gaussian mixture models (GMM) and Mahalanobis distance (MD). GMM clusters the datasets of extracted variable normally performed to form the health benchmark, and utilize MD to analyze the deviation of real time state of papermaking process from health, and determine whether to warn the operators of paper break through kernel density estimation. The verification results showed that the proposed model has a fault prediction accuracy of 76.8% and a recall rate of 72.5%, which allows paper break associated anomalous data to be observed in advance, providing valuable time for subsequent fault diagnosis.Note to Practitioners—This article is motivated by the problems of paper break in the papermaking process, which can be applied to fault prediction in papermaking and other associated complex process industries. Due to technical limitations, it is difficult to monitor all parameters of the entire production process to prevent faults from the manufacturing processes. This paper exploits existing imperfect production data, and analyzes breaking mechanisms in the papermaking process to interpret the meaning of certain data. It is obtained variables closely relating to production through data analysis and a framework consisting of health benchmark with deviation determination. Studies on corresponding actual production cases validated that the approach is feasible. The paper breaks can be predicted, but there are still some faults of which mechanisms are too unclear to support prediction. In the future research, it is encouraged to improve the accuracy and timeliness of prediction.
Zhenglei He, Guojian Chen, Mengna Hong, Qingang Xiong, Xianyi Zeng, Yi Man
IEEE Trans Autom. Sci. Eng.6
2024 DCBFusion: an infrared and visible image fusion method through detail enhancement, contrast reserve and brightness balance
Shenghui Sun, Kechen Song, Yi Man, Hongwen Dong, Yunhui Yan
Vis. Comput.3
2023 SISG-Net: Simultaneous instance segmentation and grasp detection for robot grasp in clutter
Yunhui Yan, Ling Tong 0006, Kechen Song, Hongkun Tian, Yi Man, Wenkang Yang
Adv. Eng. Informatics5
2018 Joint Content Caching and Delivery Policy for Heterogeneous Cellular Networks
abstract
Caching the popular contents near users is an effective way to release the burden of the wireless networks and reduce the energy consumption of content service for delivery. In this paper, we consider a distributed way to cache content with different user preference in heterogeneous cellular networks (HetNets). Aiming at minimizing the energy consumption of the whole network, a joint content cache and delivery optimization problem is proposed. Considering the coupling multiplicative variables, we utilize the alternative optimization (AO) algorithm to decompose the original problem into the content cache and delivery problems, which are solved separately through the knapsack solution and message passing (MP) algorithm. Numerical results reveal that the proposed scheme achieves more in energy saving than conventional schemes.
Yinglei Teng, Yi Man
PIMRC4
2015 A cooperative diversity transmission scheme by superposition coding relaying for a wireless system with multiple relays
Yang Liu 0024, Yi Man, Hongtao Zhang 0001, Li Wang 0039
Wirel. Networks2
2006 A Comparative Study on Text Clustering Methods
Xiaochun Cheng, Ronghuai Huang, Yi Man
ADMA4
2004 Neural network fusion strategies for identifying breast masses
abstract
In this work, we introduce the perceptron average neural network fusion strategy and implemented a number of other fusion strategies to identify breast masses in mammograms as malignant or benign with both balanced and imbalanced input features. We numerically compare various fixed and trained fusion rules, i.e., the majority vote, simple average, weighted average, and perceptron average, when applying them to a binary statistical pattern recognition problem. To judge from the experimental results, the weighted average approach outperforms the other fusion strategies with balanced input features, while the perceptron average is superior and achieves the goals with lowest standard deviation with imbalanced ensembles. We concretely analyze the results of above fusion strategies, state the advantages of fusing the component networks, and provide our particular broad sense perspective about information fusion in neural networks.
Yi Man, Juan Ignacio Arribas
IJCNN3