Shan Tang

dblp:87/661 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 73% Performance modeling and evaluation · 16% Processor architecture and microarchitecture · 11%
Software engineering, system software, and programming languages
2 papers
Software maintenance and evolution · 46% Empirical software engineering · 28% Compilers and program optimization · 26%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution › code recommendation
API recommendation
0.922021
Speeding Up Data Manipulation Tasks with Alternative Implementations: An Exploratory Study · ACM Trans. Softw. Eng. Methodol. 2021
How Do API Selections Affect the Runtime Performance of Data Analytics Tasks? · ASE 2019
Hardware security and side channels
fault attacks
0.912025
ρHammer: Reviving RowHammer Attacks on New Architectures via Prefetching · MICRO 2025
Hardware security and side channels › fault attacks › fault injection attack
rowhammer attack
0.912025
ρHammer: Reviving RowHammer Attacks on New Architectures via Prefetching · MICRO 2025
Memory systems
DRAM
0.912025
ρHammer: Reviving RowHammer Attacks on New Architectures via Prefetching · MICRO 2025
Memory systems › memory controller
DRAM address mapping
0.912025
ρHammer: Reviving RowHammer Attacks on New Architectures via Prefetching · MICRO 2025
Empirical software engineering
mining software repositories
0.322021
Speeding Up Data Manipulation Tasks with Alternative Implementations: An Exploratory Study · ACM Trans. Softw. Eng. Methodol. 2021
How Do API Selections Affect the Runtime Performance of Data Analytics Tasks? · ASE 2019
Empirical software engineering › mining software repositories › question and answer sites
stack overflow analysis
0.322021
Speeding Up Data Manipulation Tasks with Alternative Implementations: An Exploratory Study · ACM Trans. Softw. Eng. Methodol. 2021
How Do API Selections Affect the Runtime Performance of Data Analytics Tasks? · ASE 2019
Processor architecture and microarchitecture
speculative execution
0.312025
ρHammer: Reviving RowHammer Attacks on New Architectures via Prefetching · MICRO 2025

Methods — techniques the papers use, named apart from their topics

pseudo-barriers · 1.7prefetch-based hammering · 1.7multi-bank parallelism · 1.7control-flow obfuscation · 0.9control flow obfuscation · 0.9comparative structure mining · 0.8empirical study · 0.5comparative structure extraction · 0.5
YearPublicationVenuePosition
2025 ρHammer: Reviving RowHammer Attacks on New Architectures via Prefetching
abstract
Rowhammer is a critical vulnerability in dynamic random access memory (DRAM) that continues to pose a significant threat to various systems. However, we find that conventional load-based attacks are becoming highly ineffective on the most recent architectures such as Intel Alder and Raptor Lake. In this paper, we present $ρ$Hammer, a new Rowhammer framework that systematically overcomes three core challenges impeding attacks on these new architectures. First, we design an efficient and generic DRAM address mapping reverse-engineering method that uses selective pairwise measurements and structured deduction, enabling recovery of complex mappings within seconds on the latest memory controllers. Second, to break through the activation rate bottleneck of load-based hammering, we introduce a novel prefetch-based hammering paradigm that leverages the asynchronous nature of x86 prefetch instructions and is further enhanced by multi-bank parallelism to maximize throughput. Third, recognizing that speculative execution causes more severe disorder issues for prefetching, which cannot be simply mitigated by memory barriers, we develop a counter-speculation hammering technique using control-flow obfuscation and optimized NOP-based pseudo-barriers to maintain prefetch order with minimal overhead. Evaluations across four latest Intel architectures demonstrate $ρ$Hammer's breakthrough effectiveness: it induces up to 200K+ additional bit flips within 2-hour attack pattern fuzzing processes and has a 112x higher flip rate than the load-based hammering baselines on Comet and Rocket Lake. Also, we are the first to revive Rowhammer attacks on the latest Raptor Lake architecture, where baselines completely fail, achieving stable flip rates of 2,291/min and fast end-to-end exploitation.
Shan Tang, Yulin Tang, Xiapu Luo, Yinqian Zhang, Weizhong Qiang
MICRO2
2025 Knowledge distillation via teacher-modeled sample relationships for skin cancer diagnosis
Peng Liu 0056, Wenhua Qian, Shan Tang
Eng. Appl. Artif. Intell.3
2025 Art creator: Steering styles in diffusion model
Shan Tang, Wenhua Qian, Peng Liu 0056, Jinde Cao
Neurocomputing1
2024 Synthetic lethal connectivity and graph transformer improve synthetic lethality prediction
abstract
Synthetic lethality (SL) has shown great promise for the discovery of novel targets in cancer. CRISPR double-knockout (CDKO) technologies can only screen several hundred genes and their combinations, but not genome-wide. Therefore, good SL prediction models are highly needed for genes and gene pairs selection in CDKO experiments. However, lack of scalable SL properties prevents generalizability of SL interactions to out-of-sample data, thereby hindering modeling efforts. In this paper, we recognize that SL connectivity is a scalable and generalizable SL property. We develop a novel two-step multilayer encoder for individual sample-specific SL prediction model (MLEC-iSL), which predicts SL connectivity first and SL interactions subsequently. MLEC-iSL has three encoders, namely, gene, graph, and transformer encoders. MLEC-iSL achieves high SL prediction performance in K562 (AUPR, 0.73; AUC, 0.72) and Jurkat (AUPR, 0.73; AUC, 0.71) cells, while no existing methods exceed 0.62 AUPR and AUC. The prediction performance of MLEC-iSL is validated in a CDKO experiment in 22Rv1 cells, yielding a 46.8% SL rate among 987 selected gene pairs. The screen also reveals SL dependency between apoptosis and mitosis cell death pathways.
Kunjie Fan, Birkan Gokbag, Shan Tang, Shangjia Li, Yirui Huang, Lijun Cheng
Briefings Bioinform.3
2024 STW-MD: a novel spatio-temporal weighting and multi-step decision tree method for considering spatial heterogeneity in brain gene expression data
abstract
Gene expression during brain development or abnormal development is a biological process that is highly dynamic in spatio and temporal. Previous studies have mainly focused on individual brain regions or a certain developmental stage. Our motivation is to address this gap by incorporating spatio-temporal information to gain a more complete understanding of brain development or abnormal brain development, such as Alzheimer's disease (AD), and to identify potential determinants of response. In this study, we propose a novel two-step framework based on spatial-temporal information weighting and multi-step decision trees. This framework can effectively exploit the spatial similarity and temporal dependence between different stages and different brain regions, and facilitate differential gene analysis in brain regions with high heterogeneity. We focus on two datasets: the AD dataset, which includes gene expression data from early, middle and late stages, and the brain development dataset, spanning fetal development to adulthood. Our findings highlight the advantages of the proposed framework in discovering gene classes and elucidating their impact on brain development and AD progression across diverse brain regions and stages. These findings align with existing studies and provide insights into the processes of normal and abnormal brain development.
Shanjun Mao, Runjiu Chen, Yizhu Diao, Zongjin Li, Qingzhe Wang, Shan Tang, Shuixia Guo
Briefings Bioinform.8
2023 Arbitrary style transfer based on Attention and Covariance-Matching
Haiyuan Peng, Wenhua Qian, Jinde Cao, Shan Tang
Comput. Graph.4
2021 Speeding Up Data Manipulation Tasks with Alternative Implementations: An Exploratory Study
abstract
As data volume and complexity grow at an unprecedented rate, the performance of data manipulation programs is becoming a major concern for developers. In this article, we study how alternative API choices could improve data manipulation performance while preserving task-specific input/output equivalence. We propose a lightweight approach that leverages the comparative structures in Q&A sites to extracting alternative implementations. On a large dataset of Stack Overflow posts, our approach extracts 5,080 pairs of alternative implementations that invoke different data manipulation APIs to solve the same tasks, with an accuracy of 86%. Experiments show that for 15% of the extracted pairs, the faster implementation achieved >10x speedup over its slower alternative. We also characterize 68 recurring alternative API pairs from the extraction results to understand the type of APIs that can be used alternatively. To put these findings into practice, we implement a tool, AlterApi7 , to automatically optimize real-world data manipulation programs. In the 1,267 optimization attempts on the Kaggle dataset, 76% achieved desirable performance improvements with up to orders-of-magnitude speedup. Finally, we discuss notable challenges of using alternative APIs for optimizing data manipulation programs. We hope that our study offers a new perspective on API recommendation and automatic performance optimization.
Yida Tao, Shan Tang, Yepang Liu 0001, Zhiwu Xu 0001, Shengchao Qin
ACM Trans. Softw. Eng. Methodol.2
2019 How Do API Selections Affect the Runtime Performance of Data Analytics Tasks?
abstract
As data volume and complexity grow at an unprecedented rate, the performance of data analytics programs is becoming a major concern for developers. We observed that developers sometimes use alternative data analytics APIs to improve program runtime performance while preserving functional equivalence. However, little is known on the characteristics and performance attributes of alternative data analytics APIs. In this paper, we propose a novel approach to extracting alternative implementations that invoke different data analytics APIs to solve the same tasks. A key appeal of our approach is that it exploits the comparative structures in Stack Overflow discussions to discover programming alternatives. We show that our approach is promising, as 86% of the extracted code pairs were validated as true alternative implementations. In over 20% of these pairs, the faster implementation was reported to achieve a 10x or more speedup over its slower alternative. We hope that our study offers a new perspective of API recommendation and motivates future research on optimizing data analytics programs.
Yida Tao, Shan Tang, Yepang Liu 0001, Zhiwu Xu 0001, Shengchao Qin
ASE2
2018 BenchIP: Benchmarking Intelligence Processors
Jinhua Tao, Zidong Du, Qi Guo 0001, Huiying Lan, Lei Zhang 0008, Shengyuan Zhou, Lingjie Xu, Shan Tang, Allen Rush, Willian Chen, Shaoli Liu, Yunji Chen, Tianshi Chen 0002
J. Comput. Sci. Technol.10
2014 System-level design methodology enabling fast development of baseband MP-SoC for 4G small cell base station
abstract
“Small Cell” is regarded as the solution to optimize 4G wireless networks with improved coverage and capacity and expected to be deploy in a large number. To meet performance requirements and special constraints on the cost and size, we design a heterogeneous multi-processor SoC for small cell base station, which is composed of ASP (Application Specific Processor) cores, hardware accelerators, general-purpose processor core, and infrastructure and interface blocks. The challenges of developing such a complex chip drive us to employ system-level design methodology in both single core and mutli-core architecture optimizations. The paper discusses in detail the LISA (Language for Instruction-Set Architectures)/SystemC based ASP-algorithm joint optimization, and task-graph driven multi-core architecture exploration. Finally, the results of silicon implementation on SMIC 55nm technology are presented.
Shan Tang, Yongtao Su
DATE1
2014 Towards Sustainability-Oriented Development of Dynamic Reconfigurable Software Systems
Shan Tang, Jianxin Xue
SEKE1
2013 Supporting Integration of COTS Components from a Perspective of Self-Adaptive Software Architecture
abstract
Component-Based Software Development (CBSD) provides a high efficient and low cost way to construct software systems by integrating reusable software components. Although CBSD has already become a widely accepted paradigm, it is still beyond possibility to assemble components easily from commercial-off-the-shelf (COTS) components into one application system. A variety of mismatches between components often impede the integration of COTS components, and component adaptation is becoming a key problem in Component-Based Software Engineering (CBSE). Aiming at this requirement, this paper presents a self-adaptive software architecture (model-based) approach for supporting seamless integration of COTS components. Specifically, we first propose a self-adaptive software architecture model, and then we discuss and exemplify how to eliminate the mismatches between heterogeneous COTS components based on this model. A simplified on-line shopping system is referred throughout the paper to illustrate our approach.
Shan Tang
COMPSAC1
2013 A 100 GOPS ASP based baseband processor for wireless communication
abstract
This paper presents an ASP (application specific processor) with 512-bit SIMD (Single Instruction Multiple Data) and 192-bit VLIW (Very Long Instruction Word) architecture optimized for wireless baseband processing. It employs optimized architecture and address generation unit to accelerate the kernel algorithms. Based on the ASP, a multi-core baseband processor is developed which can work at 2×2 MIMO and 20 MHz physical bandwidth configuration for LTE inner receiver and meet requirements of Category 3 User Equipment (CAT3 UE). Furthermore, a silicon implementation of the baseband processor with 130nm CMOS technology is presented. Experimental results show that the baseband processor provides 100 GOPS computing ability at 117.6MHz.
Shan Tang, Yongtao Su, Juan Han, Jinglin Shi
DATE2
2013 Simplified MMSE Detectors for Turbo Receiver in BICM MIMO Systems
Juan Han, Qiuju Wang, Shan Tang
J. Comput. Sci. Technol.5
2010 Robust Downlink Precoding in Multiuser MIMO-OFDM Systems with Time-Domain Quantized Feedback
abstract
We consider the robust linear precoding (LP) and Tomlinson-Harashima precoding (THP) schemes for multiuser MIMO-OFDM downlink channels with limited feedback. Benefiting from the correlation of spatial channels, the mobile terminal compresses and feeds back the time-domain channel vectors instead of the corresponding frequency-domain vectors to substantially reduce the feedback signalling overhead. A compression and restoration method and a codebook design for channel state information at the transmitter (CSIT) feedback are proposed in the time domain. By treating the partial CSIT as a random quantity, we develop the robust precoders to combat the truncation and quantization errors introduced in the feedback procedure. In comparison with the non-robust designs, both the robust LP and THP have better bit-error rate performance especially in high signal-to-noise ratio region.
Yongtao Su, Shan Tang, Jinglin Shi, Xiaojing Huang 0001, Y. Jay Guo
WCNC2
2008 A debug probe for concurrently debugging multiple embedded cores and inter-core transactions in NoC-based systems
abstract
Existing SoC debug techniques mainly target bus-based systems. They are not readily applicable to the emerging system that use Network-on-Chip (NoC) as on-chip communication scheme. In this paper, we present the detailed design of a novel debug probe (DP) inserted between the core under debug (CUD) and the NoC. With embedded configurable triggers, delay control and timestamping mechanism, the proposed DP is very effective for inter-core transaction analysis as well as controlling embedded cores’ debug processes. Experimental results show the functionalities of the proposed DP and its area overhead.
Shan Tang, Qiang Xu 0001
ASP-DAC1
2008 An Adaptive Software Architecture Model Based on Component-Mismatches Detection and Elimination
abstract
Commercial-off-the-shelf components (COTS) are widely reused at present and black-box composition is the unique way to integrate them into the target system. However, various mismatches among components often hamper the integration of COTS. While the existing approaches to modeling software systems often neglect the issue of COTS component-mismatch detection and elimination. Aiming at solving this problem, this paper proposes an adaptive software architecture model. Based on this model, we first analyze and conclude the mismatches among heterogeneous components, and then we propose the corresponding solutions to eliminate these mismatches and provide a seamless integration for COTS components. At the same time, we employ an example of Web-based application system to illustrate the efficiency of our approach.
Shan Tang, Xin Peng 0001, Yiming Lau, Wenyun Zhao, Zhixiong Jiang
COMPSAC1
2008 In-band Cross-Trigger Event Transmission for Transaction-Based Debug
abstract
Cross-trigger, the mechanism to trigger activities in one debug entity from debug events happened in another debug entity, is a very useful technique for debugging applications involving multiple embedded cores. Existing solutions rely on dedicated interconnects (i.e., different from functional interconnects) to transfer debug events and cannot guarantee the arrival time of the debug events coincides with the arrival time of the data messages between multiple cores. This results in mismatches between the observed system internal operations and the ones that designers expect to watch. To tackle the above problem, in this paper, we propose to package the cross-trigger events and the actual data together into transaction messages and transfer them along the same functional interconnects (namely in- band debug event transmission), with the help of novel design-for- debug circuits. Simulation results on a hypothetical NoC-based systems show the effectiveness of the proposed technique.
Shan Tang, Qiang Xu 0001
DATE1
2007 A Connector-Centric Approach to Aspect-Oriented Software Evolution
abstract
Lose sight of the existence of system crosscutting concerns, e.g. safety and quality etc, often causes the system hard to maintain and evolve according to the changing environment and requirements. In this paper we propose an incremental aspect-oriented (AO) approach to ease this kind of evolution problem in architecture level. In this approach we introduce a novel connector, namely aspect weaving connector (AWC), to support the seamless integration of AOSD and software architecture modeling. Concretely crosscutting concerns are encapsulated into aspects and modeled as software components. AWC acts as a connector wrapper coordinating the interaction between aspectual and regular components. In order to provide a formal basic to AWC, we propose a conceptual model of it, which formalizes the underlying mechanisms of aspect dynamic weaving in architecture level using process algebra CSP. Then we verify the model's properties with FDR2 and prove that our connector-centric AO architecture modeling approach can give system an architectural dynamism and make it easier to maintain and evolve.
Yiming Lau, Wenyun Zhao, Xin Peng 0001, Shan Tang
COMPSAC (2)4
2007 A multi-core debug platform for NoC-based systems
abstract
Network-on-chip (NoC) is generally regarded as the most promising solution for the future on-chip communication scheme in giga-scale integrated circuits. As traditional debug architecture for bus-based systems is not readily applicable to identify bugs in NoC-based systems, in this paper, we present a novel debug platform that supports concurrent debug access to the cores under debug (CUDs) and the NoC in a unified architecture. By introducing core-level debug probes in between the CUDs and their network interfaces and a system-level debug agent controlled by an off-chip multi-core debug controller, the proposed debug platform provides in-depth analysis features for NoC-based systems, such as NoC transaction analysis, multi-core cross-triggering and global synchronized timestamping. Therefore, the proposed solution is expected to facilitate the designers to identify bugs in NoC-based systems more effectively and efficiently. Experimental results show that the design-for-debug cost for the proposed technique in terms of area and traffic requirements is moderate
Shan Tang, Qiang Xu 0001
DATE1