Ruiming Chen

dblp:34/6883 · DBLP profile ↗
← Back
17ranked-venue papers
15as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 13 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 33% Efficient and distributed learning · 25% Deep learning architectures and training · 25%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Electronic design automation · 90% Integrated circuit design · 6% Performance modeling and evaluation · 4%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
CLIP
1.012026
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge · AAAI 2026
Computer vision › Vision and language
vision-language pretraining
1.012026
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
0.812024
Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
model initialization
0.812024
Transformer as Linear Expansion of Learngene · AAAI 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
Transformer as Linear Expansion of Learngene · AAAI 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale Models · NeurIPS 2024
Electronic design automation
physical design
0.232007
An Effective Algorithm for Buffer Insertion in General Circuits Based on Network Flow · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Fast Min-Cost Buffer Insertion under Process Variations · DAC 2007
An Efficient Data Structure for Maxplus Merge in Dynamic Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › physical design
buffer insertion
0.122007
An Effective Algorithm for Buffer Insertion in General Circuits Based on Network Flow · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Fast Min-Cost Buffer Insertion under Process Variations · DAC 2007
Electronic design automation › physical design › buffer insertion
slew-constrained buffering
0.112007
Fast Min-Cost Buffer Insertion under Process Variations · DAC 2007
Electronic design automation › physical design
timing optimization
0.112007
Fast Min-Cost Buffer Insertion under Process Variations · DAC 2007
Integrated circuit design
clocking
0.112006
Statistical timing verification for transparently latched circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › physical design
floorplanning
0.112006
An Efficient Data Structure for Maxplus Merge in Dynamic Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation
logic synthesis
0.112006
An Efficient Data Structure for Maxplus Merge in Dynamic Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › physical design › floorplanning
slicing floorplan
0.112006
An Efficient Data Structure for Maxplus Merge in Dynamic Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › hardware verification and test › timing verification
statistical timing verification
0.112006
Statistical timing verification for transparently latched circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › logic synthesis
technology mapping
0.112006
An Efficient Data Structure for Maxplus Merge in Dynamic Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › hardware verification and test
timing verification
0.112006
Statistical timing verification for transparently latched circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Performance modeling and evaluation › simulation
monte carlo simulation
0.012006
Statistical timing verification for transparently latched circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Performance modeling and evaluation
simulation
0.012006
Statistical timing verification for transparently latched circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006

Methods — techniques the papers use, named apart from their topics

learngene · 1.5weighted-sum initialization · 1.0multimodal block extraction · 1.0soft distillation · 0.8linear expansion · 0.8linear decomposition · 0.8dynamic programming · 0.1network flow · 0.1convex-cost-flow · 0.1positive cycle probability computation · 0.1graph traversal · 0.1cycle-breaking heuristic · 0.1balanced binary search tree · 0.1
YearPublicationVenuePosition
2026 Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
abstract
CLIP (Contrastive Language-Image Pre-training) has attracted widespread attention for its multimodal generalizable knowledge, which is significant for downstream tasks. However, the computational overhead of a large number of parameters and large-scale pre-training poses challenges of pre-training a different scale of CLIP. Learngene extracts the generalizable components termed as learngene from an ancestry model and initializes diverse descendant models with it. Previous Learngene paradigms fail to handle the generalizable knowledge in multimodal scenarios. In this paper, we put forward the idea of utilizing a multimodal block to extract the multimodal generalizable knowledge, which inspires us to propose MM-LG (Multimodal Learngene), a novel framework designed to extract and leverage generalizable components from CLIP. Specifically, we first establish multimodal and unimodal blocks to extract the multimodal and unimodal generalizable knowledge in a weighted-sum manner. Subsequently, we employ these components to numerically initialize descendant models of varying scales and modalities. Extensive experiments demonstrate MM-LG's effectiveness, which achieves performance gains over existing learngene approaches (e.g.,+3.1% on Oxford-IIIT PET and +4.13% on Flickr30k) and comparable or superior results to the pre-training and fine-tuning paradigm (e.g.,+1.9% on Oxford-IIIT PET and +3.65% on Flickr30k). Notably, MM-LG requires only around 25% of the parameter storage while reducing around 2.8× pre-training costs for diverse model scales compared to the pre-training and fine-tuning paradigm, making it particularly suitable for efficient deployment across diverse downstream tasks.
Ruiming Chen, Shiyu Xia, Xu Yang 0021, Xin Geng 0001
AAAI1
2024 Transformer as Linear Expansion of Learngene
abstract
We propose expanding the shared Transformer module to produce and initialize Transformers of varying depths, enabling adaptation to diverse resource constraints. Drawing an analogy to genetic expansibility, we term such module as learngene. To identify the expansion mechanism, we delve into the relationship between the layer's position and its corresponding weight value, and find that linear function appropriately approximates this relationship. Building on this insight, we present Transformer as Linear Expansion of learnGene (TLEG), a novel approach for flexibly producing and initializing Transformers of diverse depths. Specifically, to learn learngene, we firstly construct an auxiliary Transformer linearly expanded from learngene, after which we train it through employing soft distillation. Subsequently, we can produce and initialize Transformers of varying depths via linearly expanding the well-trained learngene, thereby supporting diverse downstream scenarios. Extensive experiments on ImageNet-1K demonstrate that TLEG achieves comparable or better performance in contrast to many individual models trained from scratch, while reducing around 2× training cost. When transferring to several downstream classification datasets, TLEG surpasses existing initialization methods by a large margin (e.g., +6.87% on iNat 2019 and +7.66% on CIFAR-100). Under the situation where we need to produce models of varying depths adapting for different resource constraints, TLEG achieves comparable results while reducing around 19× parameters stored to initialize these models and around 5× pre-training costs, in contrast to the pre-training and fine-tuning approach. When transferring a fixed set of parameters to initialize different models, TLEG presents better flexibility and competitive performance while reducing around 2.9× parameters stored to initialize, compared to the pre-training approach.
Shiyu Xia, Miaosen Zhang, Xu Yang 0021, Ruiming Chen, Xin Geng 0001
AAAI4
2024 Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale Models
abstract
Vision Transformers (ViTs) are widely used in a variety of applications, while they usually have a fixed architecture that may not match the varying computational resources of different deployment environments. Thus, it is necessary to adapt ViT architectures to devices with diverse computational overheads to achieve an accuracy-efficient trade-off. This concept is consistent with the motivation behind Learngene. To achieve this, inspired by polynomial decomposition in calculus, where a function can be approximated by linearly combining several basic components, we propose to linearly decompose the ViT model into a set of components called learngenes during element-wise training. These learngenes can then be recomposed into differently scaled, pre-initialized models to satisfy different computational resource constraints. Such a decomposition-recomposition strategy provides an economical and flexible approach to generating different scales of ViT models for different deployment scenarios. Compared to model compression or training from scratch, which require to repeatedly train on large datasets for diverse-scale models, such strategy reduces computational costs since it only requires to train on large datasets once. Extensive experiments are used to validate the effectiveness of our method: ViTs can be decomposed and the decomposed learngenes can be recomposed into diverse-scale ViTs, which can achieve comparable or better performance compared to traditional model compression and pre-training methods. The code for our experiments is available in the supplemental material.
Shuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang 0021, Qiufeng Wang 0002, Xin Geng 0001
NeurIPS3
2008 Static timing: Back to our roots
abstract
Existing static timing methodologies apply various techniques to address increasingly larger process variations. The techniques include multi-corner timing, on-chip variation (OCV) derating coefficients, and path-based common path pessimism removal (CPPR) procedures. These techniques, however, destroy the benefits of linear run-time and incrementality possessed by classical static timing. The major contribution of this work is an efficient statistical timing methodology with comprehensive modeling of process variations, while at the same time retaining those key benefits. Our methodology is compatible with existing characterization methods and scales well to large chip designs. To achieve this goal, three techniques are developed: (1) building the statistical delay model based on existing multi-corner library characterization; (2) modeling spatial correlation in a scalable manner; and (3) avoiding the time-consuming CPPR procedure by removing common path pessimism in the clock network by an incremental block-based technique. Experimental results on industrial 90 nm ASIC designs show that the proposed timing methodology correctly handles all types of process variation, achieves high correlation with traditional multi-corner timing with more than 4 × speedup, and is a vehicle for pessimism reduction.
Ruiming Chen, Lizheng Zhang, Vladimir Zolotov, Chandu Visweswariah, Jinjun Xiong
ASP-DAC1
2008 Fast Estimation of Timing Yield Bounds for Process Variations
abstract
With aggressive scaling down of feature sizes in VLSI fabrication, process variation has become a critical issue in designs. We show that two necessary conditions for the ldquomaxrdquo operation are actually not satisfied in the moment matching based statistical timing analysis approaches. We propose two correlation-aware block-based statistical timing analysis approaches that keep these necessary conditions, and show that our approaches always achieve the lower bound and the upper bound on the timing yield. Our approach combining with moment-matching based statistical static timing analysis (SSTA) approaches can efficiently estimate the maximal possible errors of moment-matching-based SSTA approaches.
Ruiming Chen, Hai Zhou 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Fast Buffer Insertion for Yield Optimization Under Process Variations
abstract
With the emerging process variations in fabrication, the traditional corner-based timing optimization techniques become prohibitive. Buffer insertion is a very useful technique for timing optimization. In this paper, we propose a buffer insertion algorithm with the consideration of process variations. We use the solutions from the deterministic buffering that sets all the random variables at their nominal values to guide the statistical buffering algorithm. Our algorithm keeps the solution lists short, and always achieves higher yield than the deterministic buffering. The experimental results demonstrate that the exiting approaches cannot handle large cases efficiently or effectively, while our algorithm handles large cases very efficiently, and improves the yield more than 12% on average.
Ruiming Chen, Hai Zhou 0001
ASP-DAC1
2007 New Block-Based Statistical Timing Analysis Approaches Without Moment Matching
abstract
With aggressive scaling down of feature sizes in VLSI fabrication, process variation has become a critical issue in designs. We show that two necessary conditions for the "Max" operation are actually not satisfied in the moment matching based statistical timing analysis approaches. We propose two correlation-aware block-based statistical timing analysis approaches that keep these necessary conditions, and prove that our approaches always achieve tight lower bound and upper bound of the yield. Especially, our approach always gets the tight upper bound of the yield irrespective of the distributions that random variables have.
Ruiming Chen, Hai Zhou 0001
ASP-DAC1
2007 Fast Min-Cost Buffer Insertion under Process Variations
abstract
Process variation has become a critical problem in modern VLSI fabrication. In the presence of process variation, buffer insertion problem under performance constraints becomes more difficult since the solution space expands greatly. We propose efficient dynamic programming approaches to handle the min-cost buffer insertion under process variations. Our approaches handle delay constraints and slew constraints, in trees and in combinational circuits. The experimental results demonstrate that in general, process variations have great impact on slew-constrained buffering, but much less impact on delay-constrained buffering, especially for small nets. Our approaches have less than 9% runtime overhead on average compared with a single pass of deterministic buffering for delay constrained buffering, and get 56% yield improvement and 11.8% buffer area reduction, on average, for slew constrained buffering.
Ruiming Chen, Hai Zhou 0001
DAC1
2007 Timing budgeting under arbitrary process variations
abstract
Timing budgeting under process variations is an important step in a statistical optimization flow. We propose a novel for- mulation of the problem where budgets are statistical instead of deterministic as in existing works. This new formulation considers the changes of both the means and variances of de- lays, and thus can reduce the timing violation introduced by ignoring the changes of variances. We transform the problem to a linear programming problem using a robust optimization technique. Our approach can be used in late-stage design where the detailed distribution information is known, and is most useful in early-stage design since our approach does not assume specific underlying distributions. In addition, with the help of block-level timing budgeting, our approach can reduce the timing pessimism. Our approach is applied to the leakage power minimization problem. The results demon- strate that our approach can reduce timing violation from 690ps to 0ps, and the worst total leakage power by 17.50% on average.
Ruiming Chen, Hai Zhou 0001
ICCAD1
2007 An Effective Algorithm for Buffer Insertion in General Circuits Based on Network Flow
abstract
The problem of buffer insertion in a single net has been the focus of most previous research works. However, effective algorithms for buffer insertion in whole circuits are generally needed. In this paper, we relate the timing-constrained minimal buffer insertion problem to the convex cost-flow dual problem and propose an algorithm based on the convex cost-flow theory to solve it in combinational circuits. Experimental results demonstrate that our approach is effective. On the average, for the cases where buffering locations are not specified, our approach achieves a 46% reduction on the total buffer area in comparison to a traditional approach; for the cases where buffering locations are specified, our approach achieves a 52% reduction.
Ruiming Chen, Hai Zhou 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 Statistical timing verification for transparently latched circuits
abstract
High-performance integrated-circuit designs need to verify the clock schedules as they usually have level-sensitive latches for their speed. With process variations, the verification needs to compute the probability of correct clocking. Because of complex statistical correlations and accumulated inaccuracy of statistical operations, traditional iterative approaches have difficulties in getting accurate results. A statistical check of the structural conditions for correct clocking is proposed instead, where the central problem is to compute the probability of having a positive cycle in a graph with random edge weights. The authors proposed two algorithms to handle this. The proposed algorithms traverse the graph only several times to reduce the correlations among iterations, and it considers not only data delay variations but also clock-skew variations. Although the first algorithm is a heuristic algorithm that may overestimate timing yields, experimental results show that it has an error of 0.16% on average in comparison with the Monte Carlo (MC) simulation. Based on a cycle-breaking technique, the second heuristic algorithm can conservatively estimate timing yields. Both algorithms are much more efficient than the MC simulation.
Ruiming Chen, Hai Zhou 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 An Efficient Data Structure for Maxplus Merge in Dynamic Programming
abstract
Dynamic programming is a useful technique to handle slicing floorplan, technology mapping, and buffering problems, where many maxplus merge operations of solution lists are needed. Shi proposed an efficient O(nlogn) time algorithm to speed up the merge operation. Based on balanced binary search trees, his algorithm showed superb performance with the most unbalanced sizes of merging solution lists. The authors propose in this paper a more efficient data structure for the merge operations. With parameters to adjust adaptively, their algorithm works better than Shi's under all cases, unbalanced, balanced, and mix sizes. Their data structure is also simpler
Ruiming Chen, Hai Zhou 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2005 Efficient algorithms for buffer insertion in general circuits based on network flow
abstract
With shrinking VLSI feature sizes and increasing overall chip areas, buffering has emerged as an effective solution to the problem of growing interconnect delays in modern designs. The problem of buffer insertion in a single net has been the focus of most previous researches. However, efficient algorithms for buffer insertion in whole circuits are generally needed. In this paper, we relate the timing constrained minimal buffer insertion problem to the min-cost flow dual problem, and propose two algorithms based on min-cost flow and min-cut techniques, respectively, to solve it in combinational circuits. We compare our approaches to a traditional approach based on Lagrangian relaxation. Experimental results demonstrate that our approaches are efficient and effective. On the average, our approaches achieve 45% and 39% reduction, respectively, on the number of buffers inserted in comparison to the traditional approach.
Ruiming Chen, Hai Zhou 0001
ICCAD1
2004 Timing macro-modeling of IP blocks with crosstalk
abstract
With the increase of design complexities and the decrease of minimal feature sizes, IP reuse is becoming a common practice while crosstalk is becoming a critical issue that must be considered. This work presents two macro-models for specifying the timing behaviors of combinational hard IP blocks with crosstalk effects. The gray-box model keeps a coupling graph and lists the conditions on relative input arrival time combinations for couplings not to take effect. The black-box model stores the output response windows for a basic set of relative input arrival time combinations, and computes the output arrival time for any given input arrival time combination through the union of some combinations in the basic set. Both macro-models are conservative, and can greatly reduce the pessimism existing in the conventional "pin-to-pin" model. This is the first work to deal with timing macro-modeling of combinational hard IP blocks with the consideration of crosstalk effects.
Ruiming Chen, Hai Zhou 0001
ICCAD1
2004 Clock schedule verification under process variations
abstract
With aggressive scaling down of feature sizes in VLSI fabrication, process variations have become a critical issue in designs, especially for high-performance ICs. Usually having level-sensitive latches for their speed, high-performance IC designs need to verify the clock schedules. With process variations, the verification needs to compute the probability of correct clocking. Because of complex statistical correlations, traditional iterative approaches are difficult to get accurate results. Instead, a statistical checking of the structural conditions for correct clocking is proposed, where the central problem is to compute the probability of having a positive cycle in a graph with random edge weights. The proposed method only traverses the graph once to avoid the correlations among iterations, and it considers not only data delay variations but also clock skew variations. Experimental results showed that the proposed approach has an error of 0.14% on average in comparisons with the Monte Carlo simulations.
Ruiming Chen, Hai Zhou 0001
ICCAD1
2004 A Flexible Data Structure for Efficient Buffer Insertion
abstract
With continuous down-scaling of minimum feature sizes and increasing of chip areas, buffering has become a necessary technique to control the interconnect delays in VLSI chips. Recently, Shi and Li proposed an efficient O(n log n) time algorithm to speed up buffering. Based on balanced binary search trees, their algorithm showed superb performance with the most unbalanced sizes of merging solution lists. We propose in this paper a more flexible data structure for the same buffering operations. With parameters to adjust, our algorithm works better than Shi and Li under all cases: unbalanced, balanced, and mix sizes. Our data structure is also simpler than theirs.
Ruiming Chen, Hai Zhou 0001
ICCD1
1993 Nonlinear estimation of scene parameters from digital images using zero-hit-length statistics
abstract
A zero-hit run-length probability model for image statistics is derived. The statistics are based on the lengths of runs of pixels that do not include any part of objects that define a scene model. The statistics are used to estimate the density and size of the discrete objects (modeled as disks) from images when the image pixel size is significant relative to the object size. Using different combinations of disk size, density, and image resolution (pixel size) in simulated images, parameter estimation may be used to investigate the essential invertibility of object size and density. Analysis of the relative errors and 95% confidence intervals indicates the accuracy and reliability of the estimates. An integrated parameter r, reveals relationships between errors and the combinations of the three basic parameters of object size, density, and pixel size. The method may be used to analyze real remotely sensed images if simplifying assumptions are relaxed to include the greater complexity found in real data.>
Ruiming Chen, David L. B. Jupp, Curtis E. Woodcock, Alan H. Strahler
IEEE Trans. Geosci. Remote. Sens.1