VLDB 2026 Research / reviewers in the wild / expert
Wei Chen 0009
dblp:c/WeiChen9
· DBLP profile ↗
55ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0002-2876-6687ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 20 since 2021Systems, architecture and hardware · 18 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 1Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attention to Threat-Relevant Objects: Reasoning Detection in Autonomous Driving via Multimodal Large Language ModelsabstractPerceiving threats is an innate human instinct. During driving, humans naturally focus their attention on objects that pose real potential risks. Motivated by this observation, we shift the focus from traditional class-based detection to a novel task termed threat-oriented reasoning detection in autonomous driving. This task aims to localize threat objects and reason about their threat levels from a driver-centric perspective. To support this task, we build a benchmark comprising diverse corner-case scenarios, annotated by multiple experienced drivers to reflect human-aligned threat cognition. Given the reasoning demands of this task, we then explore the capabilities of multi-modal large language models (MLLMs) and introduce two methods based on whether the MLLM supports object detection: 1) For MLLMs lacking detection capability, we introduce ThreatCoT, a plug-and-play training-free method that combines chain-of-thought (CoT) with a visual expert toolchain to support step-by-step reasoning. 2) For MLLMs with detection support, we introduce ThreatReasoner, an end-to-end reinforcement learning (RL)-based method built on the GRPO algorithm, which enables per-object reasoning through a fully unsupervised reward strategy. Both quantitative and qualitative experiments show that our methods can effectively unlock the new capabilities of MLLM in threat-oriented reasoning detection. Yu-Lin He, Wei Chen 0009, Xinbiao Gan, Siqi Wang 0001, Haotian Wang 0001, Yusong Tan |
AAAI | 2 |
| 2025 | Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic DisentanglementabstractOccupancy prediction plays a pivotal role in autonomous driving (AD) due to its capabilities of fine-grained 3D perception and general object recognition. However, existing methods often incur high computational costs, which conflict with AD's real-time demand. To this end, we redirect the focus from accuracy only to both accuracy and efficiency. By conducting a head-to-head comparison of existing methods, we find it challenging to balance accuracy and efficiency. We identify a core issue for this challenge: the strong coupling between geometry and semantics. Specifically, the predicted geometric structure (e.g., depth) guides the projection of 2D image features into 3D voxel space, which significantly affects feature discriminability and subsequent semantic learning. To address this issue, we focus on two key aspects: model design and learning strategies. 1) For model design, we propose a dual-branch network that disentangles the representation of geometry and semantics. The voxel branch utilizes a novel re-parameterized large-kernel 3D convolution to refine geometric structure efficiently, while the BEV branch employs temporal fusion and BEV encoding for efficient semantic learning. 2) For learning strategies, we propose to separate geometric learning from semantic learning by the mixup of ground-truth and predicted depths. Our method achieves 39.4% mIoU at 20 FPS on Occ3D-nuScenes, showcasing a state-of-the-art balance between accuracy and efficiency. Yu-Lin He, Wei Chen 0009, Siqi Wang 0001, Tianci Xun, Yusong Tan |
AAAI | 2 |
| 2025 | A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time AdaptationabstractRemote Sensing Image (RSI) segmentation has made significant strides, emerging as a leading solution for interpreting remote sensing data. However, due to the substantial domain gap between different remote sensors and limited computational resources, existing RSI segmentation methods often suffer from poor generalization performance. To address these challenges, we propose a COst-effective and Whole-process Domain Adaptation solution, namely COWDA, which adapts models at both the train and test time through three key phases: 1) Source-data Domain Alignment: We employ traditional image stylization techniques to translate source images into target-style alternatives, which avoids the computationally intensive need for auxiliary neural network models. 2) Target-data Train-time Fine-tuning: We propose a joint positive and negative learning (JPNL) algorithm that adds both positive and negative samples to effectively learn domain-invariant knowledge from noisy pseudo-labeled target data. 3) Test-time Adaptation: We propose an entropy-weighted test-time adaptation strategy to update the trained model with online test samples, further enhancing its performance. Extensive experiments on two widely-used domain adaptation benchmarks for remote sensing show that COWDA improves state-of-the-art counterparts by 1.4% and 2.6% F1 scores, respectively. Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Tianhang Guo, Yuhua Tang |
ICASSP | 1 |
| 2025 | Exploiting Foundation Models for Label-Efficient Few-Shot Learning via Feature Coupling: A Case Study of cardiac CT SegmentationabstractThe scarcity of labeled data poses a significant challenge for deep learning-based medical image segmentation. To address this, this study introduces the novel Foundation Model-based Few-Shot Segmentation (FM-FSS) paradigm. FM-FSS capitalizes on the knowledge distilled from pre-trained foundation models, such as the Segment Anything Model, to enhance segmentation performance in few-shot scenarios. The paradigm designs a feature coupling module that synergizes SAM’s powerful feature extraction capabilities with nnU-Net’s self-configuration strategy, enabling accurate segmentation with minimal labeled data and optional manual prompt inputs. Extensive experiments on a publicly available cardiac CT dataset demonstrate that FM-FSS outperforms state-of-the-art segmentation models. With only 20 labeled images, our method achieves an average Dice score of 94.33% and an ASD of 1.10 mm. Moreover, FM-FSS maintains its label-efficient performance in a one-shot setup, reducing the annotation requirements by at least fourfold. The code and pre-trained models will be released upon acceptance. Wei Chen 0009, Wenjuan Zhou, Tianhang Guo, Yuhua Tang |
ICASSP | 1 |
| 2025 | Gated Cross-Attention Network for Depth CompletionabstractDepth completion is a popular research direction in the field of depth estimation. The fusion of color and depth features is the critical challenge in this task, mainly due to the asymmetry between the rich scene details in color images and the sparse pixels in depth maps. To tackle this issue, we design an efficient Gated Cross-Attention Network that propagates confidence via a gating mechanism, simultaneously extracting and refining key information in both color and depth branches to achieve local spatial feature fusion. Additionally, we incorporate a Transformer-based attention network in low-dimensional space to effectively fuse global features and increase the network’s receptive field. At the same time, we use the Ray Tune mechanism with the AsyncHyperBandScheduler and the HyperOptSearch algorithm to automatically search for the optimal number of module iterations, which also allows us to achieve performance comparable to state-of-the-art methods. We conduct experiments on both indoor and outdoor scene datasets. Our fast network ranked first among real-time methods (below 30ms and 100ms), and our accurate network ranked first among all methods on the KITTI official website at the time of submission. Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang |
ICASSP | 5 |
| 2025 | Curve-Aware Gaussian Splatting for 3D Parametric Curve ReconstructionabstractThis paper presents an end-to-end framework for reconstructing 3D parametric curves directly from multi-view edge maps. Contrasting with existing two-stage methods that follow a sequential ``edge point cloud reconstruction and parametric curve fitting'' pipeline, our one-stage approach optimizes 3D parametric curves directly from 2D edge maps, eliminating error accumulation caused by the inherent optimization gap between disconnected stages. However, parametric curves inherently lack suitability for rendering-based multi-view optimization, necessitating a complementary representation that preserves their geometric properties while enabling differentiable rendering. We propose a novel bi-directional coupling mechanism between parametric curves and edge-oriented Gaussian components. This tight correspondence formulates a curve-aware Gaussian representation, \textbf{CurveGaussian}, that enables differentiable rendering of 3D curves, allowing direct optimization guided by multi-view evidence. Furthermore, we introduce a dynamically adaptive topology optimization framework during training to refine curve structures through linearization, merging, splitting, and pruning operations. Comprehensive evaluations on the ABC dataset and real-world benchmarks demonstrate our one-stage method's superiority over two-stage alternatives, particularly in producing cleaner and more robust reconstructions. Additionally, by directly optimizing parametric curves, our method significantly reduces the parameter count during training, achieving both higher efficiency and superior performance compared to existing approaches. Zhirui Gao, Renjiao Yi, Yaqiao Dai, Xuening Zhu, Wei Chen 0009, Chenyang Zhu 0002, Kai Xu 0004 |
ICCV | 5 |
| 2025 | Self-Supervised Learning of Hybrid Part-Aware 3D Representations of 2D Gaussians and Superquadrics
Zhirui Gao, Renjiao Yi, Yuhang Huang 0006, Wei Chen 0009, Chenyang Zhu 0002, Kai Xu 0004 |
ICCV | 4 |
| 2025 | KD-MedSAM: Lightweight Knowledge Distillation of Segment Anything Model for Multi-modality Medical Image Segmentation
Wenjuan Zhou, Jiuyuan Zhu, Wei Chen 0009, Chen Li 0034, Yu-Lin He |
ICIC (28) | 3 |
| 2025 | Combating the Negative Optimization in Source-Free Domain Adaptive Medical Image Segmentation via Selective Online Self-TrainingabstractSelf-Training has proven to be a simple yet effective framework in source-free domain adaptation (SFDA) for medical image segmentation. However, prevalent domain discrepancies and privacy restrictions on source domain data can lead to the negative optimization problem. Current methods predominantly focus on rigorous pseudo-label quality assessment, yet they often overlook the specific prior knowledge inherent to medical images and the potential of online self-training to combat negative optimization. To address these gaps, this paper introduces a selective online self-training method, tackling the issue from both the pseudo-label selection and learning stages. In the selection stage, we identify a domain-invariant prior related to organ shape, and propose a class-prior guided selection mechanism to mitigate class imbalance in pseudo-labels. In the learning stage, we introduce a strategy to selectively update the teacher model based on evolutionary state feedback generated by the student model. Extensive experiments conducted on two widely used benchmarks show the effectiveness of our method, which achieves notable 2.0% and 4.1% Dice improvements. Wenjuan Zhou, Wei Chen 0009, Yu-Lin He, Chen Li 0034 |
ICME | 2 |
| 2025 | Hierarchical Neural Architecture Search for Fast and Accurate Depth Completion
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang, Yu-Lin He |
ICMR | 5 |
| 2025 | Generic Objects as Pose Probes for Few-Shot View SynthesisabstractRadiance fields, including NeRFs and 3D Gaussians, demonstrate great potential in high-fidelity rendering and scene reconstruction, while they require a substantial number of posed images as input. COLMAP is frequently employed for preprocessing to estimate poses. However, COLMAP necessitates a large number of feature matches to operate effectively, and struggles with scenes characterized by sparse features, large baselines, or few-view images. We aim to tackle few-view NeRF reconstruction using only 3 to 6 unposed scene images, freeing from COLMAP initializations. Inspired by the idea of calibration boards in traditional pose calibration, we propose a novel approach of utilizing everyday objects, commonly found in both images and real life, as “pose probes”. By initializing the probe object as a cube shape, we apply a dual-branch volume rendering optimization (object NeRF and scene NeRF) to constrain the pose optimization and jointly refine the geometry. PnP matching is used to initialize poses between images incrementally, where only a few feature matches are enough. PoseProbe achieves state-of-the-art performance in pose estimation and novel view synthesis across multiple datasets in experiments. We demonstrate its effectiveness, particularly in few-view and large-baseline scenes where COLMAP struggles. In ablations, using different objects in a scene yields comparable performance, showing that PoseProbe is robust to the choice of probe objects. Our project page is available at:https://zhirui-gao.github.io/PoseProbe.github.io/ Zhirui Gao, Renjiao Yi, Chenyang Zhu 0002, Ke Zhuang, Wei Chen 0009, Kai Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Recalling Unknowns Without Losing Precision: An Effective Solution to Large Model-Guided Open World Object DetectionabstractOpen World Object Detection (OWOD) aims to adapt object detection to an open-world environment, so as to detect unknown objects and learn knowledge incrementally. Existing OWOD methods typically leverage training sets with a relatively small number of known objects. Due to the absence of generic object knowledge, they fail to comprehensively perceive objects beyond the scope of training sets. Recent advancements in large vision models (LVMs), trained on extensive large-scale data, offer a promising opportunity to harness rich generic knowledge for the fundamental advancement of OWOD. Motivated by Segment Anything Model (SAM), a prominent LVM lauded for its exceptional ability to segment generic objects, we first demonstrate the possibility to employ SAM for OWOD and establish the very first SAM-Guided OWOD baseline solution. Subsequently, we identify and address two fundamental challenges in SAM-Guided OWOD and propose a pioneering SAM-Guided Robust Open-world Detector (SGROD) method, which can significantly improve the recall of unknown objects without losing the precision on known objects. Specifically, the two challenges in SAM-Guided OWOD include: 1) Noisy labels caused by the class-agnostic nature of SAM; 2) Precision degradation on known objects when more unknown objects are recalled. For the first problem, we propose a dynamic label assignment (DLA) method that adaptively selects confident labels from SAM during training, evidently reducing the noise impact. For the second problem, we introduce cross-layer learning (CLL) and SAM-based negative sampling (SNS), which enable SGROD to avoid precision loss by learning robust decision boundaries of objectness and classification. Experiments on public datasets show that SGROD not only improves the recall of unknown objects by a large margin (~20%), but also preserves highly-competitive precision on known objects. The program codes are available at https://github.com/harrylin-hyl/SGROD. Yu-Lin He, Wei Chen 0009, Siqi Wang 0001, Tianrui Liu 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | ChatDSE: A Zero-Shot Microarchitecture Design Space Explorer Powered by GPT4.0abstractDesign Space Exploration (DSE) aims at identifying Pareto optimal synthesis configurations. Previous works require microarchitecture samples with key labels, including power and clock cycles, to train their models. However, as the chip design space expands rapidly, the cost of sampling the design space has significantly increased, due to the growing number of samples and time-consuming Very Large Scale Integration (VLSI) implementation flow. Recent advancements in Large Language Models (LLMs) have demonstrated their remarkable power in zero-shot learning tasks, presenting an innovative strategy for accomplishing DSE. Hence, this article presents ChatDSE, a zero-shot framework for DSE that is powered by the advanced capabilities of the LLM GPT4.0. Firstly, this framework analyzes the nature of the target microarchitecture and generates a corresponding system context to provide the prior knowledge of the microarchitecture. Secondly, a proposed sampling algorithm, PriorDC, identifies the most representative samples with pseudo labels. One of these samples is chosen as a baseline, whose power and clock cycles labels are set as 1, and the remaining sample labels are obtained by chatting with GPT4.0. Finally, ChatDSE engages in a dialogue with GPT4.0 to estimate the power and clock cycles of designs within the space, ultimately identifying the Pareto optimal design set. In the DSE for the RISC-V Berkeley Out-of-Order Machine (BOOM), experimental results show that ChatDSE is capable of identifying optimal designs and accelerates the exploration process by 574 times when compared to the state-of-the-art DSE methodologies. Mingxin Tang, Wei Chen 0009, Lizhou Wu, Libo Huang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | Don't Turn a Blind Eye to Localization Noise: Localization Pseudo-label Correction and Learning for Semi-Supervised Object DetectionabstractPseudo-labeling has proven to be a simple yet effective technique for semi-supervised object detection (SSOD). However, the inevitable noise problem in pseudo-labels seriously hinders SSOD methods. Existing methods primarily focus on classification noise, while the specific and non-negligible localization noise remains not well-addressed. This paper analyzes the localization noise arising from the alternating learning and generation phases. For the generation phase, we innovatively explore the self-correction ability of models, stepping beyond the simple pseudo-label selection. We propose a localization pseudo-label correction (LPC) strategy to self-correct pseudo boxes and enhance prediction stability. In the learning phase, we propose a noisy localization loss (NLL) to enlarge the penalty of inconsistent predictions, thereby improving localization accuracy. Applied to two classic SSOD methods (Soft Teacher and Unbiased Teacher) and a recent state-of-the-art method (PseCo), our approach consistently improves accuracy across all of them. Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Ke Liang 0006, Yusong Tan, Yulan Guo |
ICME | 2 |
| 2024 | Distinguishing Textual Prompt Importance: Image-Guided Text Weighting for CLIP-Based Few-shot LearningabstractFew-Shot learning deals with learning a model capable of recognizing new classes when provided with limited labeled data. Recently, Contrastive Language-Image Pre-training (CLIP) based methods have shown significant potential in this filed. However, due to the equal treatment for textual prompts, existing CLIP-based methods fail in establishing a strong text classifier, consequently limiting their performance. These textual prompts are obtained by querying large language models (LLMs), making it unreasonable to assume uniform quality across all of them. To address this issue, this paper proposes an Image-guided Text weighting (IGTW) module, which employs the feature similarity of training images and textual prompts to guide the weighting of textual prompts. After applying our method to the recent state-of-the-art method (i.e., CaFo) and classic method (i.e., Tip-Adapter), consistent improvements are achieved across 11 few-shot learning datasets, proving the effectiveness and universality of our method. Tianci Xun, Wei Chen 0009, Yu-Lin He, Yuanming Gao, Jiuyuan Zhu |
ICME | 2 |
| 2024 | Sniffing Threatening Open-World Objects in Autonomous Driving by Open-Vocabulary ModelsabstractAutonomous driving (AD) is a typical application that requires effectively exploiting multimedia information. For AD, it is critical to ensure safety by detecting unknown objects in an open world, driving the demand for open world object detection (OWOD). However, existing OWOD methods treat generic objects beyond known classes in the train set as unknown objects and prioritize recall in evaluation. This encourages excessive false positives and endangers safety of AD. To address this issue, we restrict the definition of unknown objects to threatening objects in AD, and introduce a new evaluation protocol, which is built upon a new metric named U-ARecall, to alleviate biased evaluation caused by neglecting false positives. Under the new evaluation protocol, we re-evaluate existing OWOD methods and discover that they typically perform poorly in AD. Then, we propose a novel OWOD paradigm for AD based on fine-tuning foundational open-vocabulary models (OVMs), as they can exploit rich linguistic and visual prior knowledge for OWOD. Following this new paradigm, we propose a brand-new OWOD solution, which effectively addresses two core challenges of fine-tuning OVMs via two novel techniques: 1) the maintenance of open-world generic knowledge by a dual-branch architecture; 2) the acquisition of scenario-specific knowledge by the visual-oriented contrastive learning scheme. Besides, a dual-branch prediction fusion module is proposed to avoid post-processing and hand-crafted heuristics. Extensive experiments show that our proposed method not only surpasses classic OWOD methods in unknown object detection by a large margin (∼× U-ARecall), but also notably outperforms OVMs without fine-tuning in known object detection (∼ 20% K-mAP). Our codes are available at https://github.com/harrylin-hyl/AD-OWOD. Yu-Lin He, Siqi Wang 0001, Wei Chen 0009, Tianci Xun, Yusong Tan |
ACM Multimedia | 3 |
| 2024 | Unleashing the Class-Incremental Learning Potential of Foundation Models by Virtual Feature Generation and Replay
Tianci Xun, Yu-Lin He, Wei Chen 0009 |
PRCV (5) | 4 |
| 2024 | Segmentation prompts classification: A nnUNet-based 3D transfer learning framework with ROI tokenization and cross-task attention for esophageal cancer T-stage diagnosis
Chen Li 0034, Runyuan Wang, Wei Chen 0009 |
Expert Syst. Appl. | 4 |
| 2024 | Crots: Cross-Domain Teacher-Student Learning for Source-Free Domain Adaptive Semantic Segmentation
Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034 |
Int. J. Comput. Vis. | 2 |
| 2023 | Adaptive Scale and Spatial Aggregation for Real-Time Object DetectionabstractCutting-edge real-time detectors usually reach real-time performance by adopting lightweight architectures. The accuracy of detection may be limited by their insufficient capabilities to obtain powerful feature representation, which is a notoriously onerous task in machine vision applications. Aiming at this problem, this study proposes a method of adaptive aggregation of features at both scale and spatial levels in an anchor-free framework: 1) at the scale level, a Multi-scale Point Feature Fusion (MPFF) module has been proposed to fuse point features from multiple scales via a self-adaptive re-weighting manner; 2) at the spatial level, a Restrained Deformable Convolution (R-DCN) has been designed to focus on the most informative features in a pre-defined region while avoiding the remote feature distraction. Based on R-DCN, an Adaptive Spatial Aggregation (ASA) module has been presented to alleviate the feature misalignment problem in classification and regression tasks via their respective spatial divisions. Extensive experimental results on MS COCO indicate that Adaptive Aggregation Detector (AADet) achieves a state-of-the-art detection performance, i.e., 41.8 AP at 60 FPS. Wei Chen 0009, Yu-Lin He, Zhengfa Liang, Yulan Guo |
ICASSP | 1 |
| 2023 | Domain Generalized Fundus Image Segmentation via Dual-Level MixingabstractSingle domain generalization plays a vital role in signal processing tasks, which is capable of extracting domain-invariant knowledge from a single source domain such that the learned model can be generalizable to unseen domains. However, many existing methods require the help of auxiliary tasks, bringing extra computation cost. Aiming at this pitfall, this study proposes Dual-Level Mixing (DLM) to boost the diversity of the single source domain and enhance the generalization performance. Specifically, at the input level, we apply different image augmentations to get different variants of the same image. Patches from augmented images are spatially mixed to get a perturbed image as the input, which enlarges the scale of training data and boosts the diversity. Meanwhile, at the feature level, we characterize feature statistics as Gaussian distributions. Then, we resample the feature statistics to renormalize the latent features, which simulates the potential style variance of different domains. In summary, the proposed DLM method synergies input-level and feature-level mixing strategies, leading to enhanced generalization performance. Experimental results of cross-domain fundus image segmentation demonstrate that either input-level mixing or feature-level mixing can effectively promote the performance of domain generalization. Moreover, the collaboration of dual-level mixing strategies leads to superior or comparable performance to domain adaptation counterparts that rely on target data. Xin Luo 0009, Wei Chen 0009, Chen Li 0034, Bin Zhou 0004, Yusong Tan |
ICASSP | 2 |
| 2023 | A knowledge-based learning framework for self-supervised pre-training towards enhanced recognition of biomedical microscopy imagesabstractSelf-supervised pre-training has become the priory choice to establish reliable neural networks for automated recognition of massive biomedical microscopy images, which are routinely annotation-free, without semantics, and without guarantee of quality. Note that this paradigm is still at its infancy and limited by closely related open issues: (1) how to learn robust representations in an unsupervised manner from unlabeled biomedical microscopy images of low diversity in samples? and (2) how to obtain the most significant representations demanded by a high-quality segmentation? Aiming at these issues, this study proposes a knowledge-based learning framework (TOWER) towards enhanced recognition of biomedical microscopy images, which works in three phases by synergizing contrastive learning and generative learning methods: (1) Sample Space Diversification: Reconstructive proxy tasks have been enabled to embed a priori knowledge with context highlighted to diversify the expanded sample space; (2) Enhanced Representation Learning: Informative noise-contrastive estimation loss regularizes the encoder to enhance representation learning of annotation-free images; (3) Correlated Optimization: Optimization operations in pre-training the encoder and the decoder have been correlated via image restoration from proxy tasks, targeting the need for semantic segmentation. Experiments have been conducted on public datasets of biomedical microscopy images against the state-of-the-art counterparts (e.g., SimCLR and BYOL), and results demonstrate that: TOWER statistically excels in all self-supervised methods, achieving a Dice improvement of 1.38 percentage points over SimCLR. TOWER also has potential in multi-modality medical image analysis and enables label-efficient semi-supervised learning, e.g., reducing the annotation cost by up to 99% in pathological classification. Wei Chen 0009, Chen Li 0034, Dan Chen 0001, Xin Luo 0009 |
Neural Networks | 1 |
| 2023 | Adversarial style discrepancy minimization for unsupervised domain adaptationabstractMainstream unsupervised domain adaptation (UDA) methods align feature distributions across different domains via adversarial learning. However, most of them focus on global distribution alignment, ignoring the fine-grained domain discrepancy. Besides, they generally require auxiliary models, bringing extra computation costs. To tackle these issues, this study proposes an UDA method that differentiates individual samples without the help of extra models. To this end, we introduce a novel discrepancy metric, termed style discrepancy, to distinguish different target samples. We also propose a paradigm for adversarial style discrepancy minimization (ASDM). Specifically, we fix the parameters of the feature extractor and maximize style discrepancy to update the classifier, which helps detect more hard samples. Adversely, we fix the parameters of the classifier and minimize the style discrepancy to update the feature extractor, pushing those hard samples near the support of the source distribution. Such adversary helps to progressively detect and adapt more hard samples, leading to fine-grained domain adaptation. Experiments on different UDA tasks validate the effectiveness of ASDM. Overall, without any extra models, ASDM reaches a 46.9% mIoU in the GTA5 to Cityscapes benchmark and an 84.7% accuracy in the VisDA-2017 benchmark, outperforming many existing adversarial-learning-based methods. Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034, Yusong Tan |
Neural Networks | 2 |
| 2022 | Adaptive Pseudo Labeling for Source-Free Domain Adaptation in Medical Image SegmentationabstractDomain adaptation is common but challenging in signal processing tasks due to the intrinsic discrepancy, especially in difficult-to-label medical image segmentation application scenarios. Pseudo labeling methods are widely utilized to compensate for the scarcity of annotation. However, most existing methods set the fixed thresholds to select highly-confident predictions as pseudo labels, inevitably generating false labels with noise. In this paper, we combine the dual-classifiers consistency and predictive category-aware confidence to form a novel regularization for pseudo-label denoising. The dual-classifiers consistency helps promote the robustness of pseudo labels. Meanwhile, category-aware confidence is utilized as adaptive pixel-wise weights, avoiding the need for handcrafted thresholds. The adapted model is refined by the rectified pseudo labels without source domain samples. The proposed method is model-independent and thus can be plug-and-play to improve existing UDA methods. We validated it on the cross-modality medical image segmentation and obtained more competitive results. Chen Li 0034, Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Yusong Tan |
ICASSP | 2 |
| 2021 | Tri-Directional Tasks Complementary Learning for Unsupervised Domain Adaptation of Cross-modality Medical Image Semantic SegmentationabstractCross-modality adaptation is challenging due to the internal domain discrepancy in appearance and representation. When the trained model of source domain is transferred to the target domain, the domain shift will reduce accuracy. Meanwhile, unsupervised domain adaptation has the potential to recover this degradation among medical images of different modalities, so it is of clinical significance and meaningful in bioinformatics. However, previous related works usually try to align domains in a single direction or two directions, failing to take advantage of the complementary relationship between different directions and alignment tasks. In this paper, we propose the Tri-directional learning framework to solve domain shift in the task of medical image semantic segmentation. The proposed framework is able to synergize image style transformation, mask segmentation and edge segmentation. The above three tasks are mutually boosted through complementary training in each iteration. In this way, our method performs cross-modality medical images semantic segmentation from labeled source domain (MRI) to unlabeled target domain (CT). The experimental results demonstrate the effectiveness of the proposed method. For the task of cardiac structure segmentation from cross-modality medical images, our proposed framework achieves state-of-the-art performance. The code is available at https://github.com/lichen14/TriDL. Chen Li 0034, Wei Chen 0009, Mingfei Wu, Xin Luo 0009, Yu-Lin He, Yusong Tan |
BIBM | 2 |
| 2021 | AttENT: Domain-Adaptive Medical Image Segmentation via Attention-Aware Translation and Adversarial Entropy MinimizationabstractDue to the intrinsic domain shift among different modalities, it is nontrivial to directly apply a well-trained model into other cross-modality medical images. Unsupervised domain adaptation (UDA) has the potential to reduce such domain shift. However, existing UDA methods try to align domains in either image level or in feature level, failing to consider the unified relationship between cross-modality images and their corresponding features. In this paper, we propose a novel UDA framework for domain adaptive medical image segmentation. The proposed framework synergizes both pixel space and entropy space for domain alignment. Specifically, in the pixel space, we introduce the attention mechanism into CycleGAN, and enhance the semantic and geometric consistency of the target organs during the image style transformation. In entropy space, we utilize entropy minimization principle to force consistent image segmentation between well-annotated source domain and non-annotated target domain. The aligned ensemble of two representation spaces enables a well-trained segmentation model to effectively transfer from source domain to target domain. The experimental results demonstrate the effectiveness of the proposed method. For the task of multi-organs segmentation from cross-modality medical images, our proposed framework achieves state-of-the-art performance, with some specific metric even superior to those of supervised methods. The code is available at https://github.com/lichen14/AttENT. Chen Li 0034, Xin Luo 0009, Wei Chen 0009, Yu-Lin He, Mingfei Wu, Yusong Tan |
BIBM | 3 |
| 2021 | Multi-Scale Cascade Disparity Refinement Stereo NetworkabstractStereo matching has attracted much attention in recent years. Traditional methods can quickly generate a disparity result, but the accuracy is low. On the contrary, methods based on neural networks can achieve a high accuracy level, but they are difficult to reach the real-time level. Therefore, this paper presents MCDRNet, which combines traditional methods with neural networks to achieve real-time and accurate stereo matching results. Concretely, our network first generates a rough disparity map based on the traditional ADCensus algorithm. Then we design a novel Multi-Scale Cascade Network to refine the disparity map from coarse to fine. We evaluate our best-trained model on the KITTI official website. The results show that our network is much faster than most current top-performing methods(31×than CSPN, 56×than GANet, etc.). Meanwhile, it is more accurate than traditional stereo methods(SGM, SPS-St) and other fast 2D convolution networks(Fast DS-CS, DispNetC, etc.), demonstrating the rationalities and feasibilities of our method. Xiaogang Jia, Wei Chen 0009, Zhengfa Liang, Xin Luo 0009, Mingfei Wu, Yusong Tan, Libo Huang 0002 |
ICASSP | 2 |
| 2021 | Multi-Scale Cost Volumes Cascade Network for Stereo MatchingabstractStereo matching is essential for robot navigation. However, the accuracy of current widely used traditional methods is low, while methods based on CNN need expensive computational cost and running time. This is because different cost volumes play a crucial role in balancing speed and accuracy. Thus we propose MSCVNet, which combines traditional methods and neural networks to improve the quality of cost volume. Concretely, our network first generates multiple 3D cost volumes with different resolutions and then uses 2D convolutions to construct a novel cascade hourglass network for cost aggregation. Meanwhile, we design an algorithm to distinguish and calculate the loss for discontinuous areas of the disparity result. According to the KITTI official website, our network is much faster than most top-performing methods (24than CSPN, 44than GANet, etc.). Meanwhile, compared to traditional methods (SPS-St, SGM) and other real-time stereo matching networks (Fast DS-CS, DispNetC, and RTSNet, etc.), our network achieves a big improvement in accuracy, demonstrating the feasibility and capability of the proposed method. Xiaogang Jia, Wei Chen 0009, Chen Li 0034, Zhengfa Liang, Mingfei Wu, Yusong Tan, Libo Huang 0002 |
ICRA | 2 |
| 2021 | Fast and Accurate Lane Detection via Frequency Domain LearningabstractIt is desirable to maintain both high accuracy and runtime efficiency in lane detection. State-of-the-art methods mainly address the efficiency problem by direct compression of high-dimensional features. These methods usually suffer from information loss and cannot achieve satisfactory accuracy performance. To ensure the diversity of features and subsequently maintain information as much as possible, we introduce multi-frequency analysis into lane detection. Specifically, we propose a multi-spectral feature compressor (MSFC) based on two-dimensional (2D) discrete cosine transform (DCT) to compress features while preserving diversity information. We group features and associate each group with an individual frequency component, which incurs only 1/7 overhead of one-dimensional convolution operation but preserves more information. Moreover, to further enhance the discriminability of features, we design a multi-spectral lane feature aggregator (MSFA) based on one-dimensional (1D) DCT to aggregate features from each lane according to their corresponding frequency components. The proposed method outperforms the state-of-the-art methods (including LaneATT and UFLD) on TuSimple, CULane, and LLAMAS benchmarks. For example, our method achieves 76.32% F1 at 237 FPS and 76.98% F1 at 164 FPS on CULane, which is 1.23% and 0.30% higher than LaneATT. Our code and models are available at https://github.com/harrylin-hyl/MSLD. Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Dan Chen 0001, Yusong Tan, Xin Luo 0009, Chen Li 0034, Yulan Guo |
ACM Multimedia | 2 |
| 2021 | Toward security as a service: A trusted cloud service architecture with policy customization
Chenlin Huang, Wei Chen 0009, Songlei Jian, Yusong Tan, Dan Chen 0001 |
J. Parallel Distributed Comput. | 2 |
| 2021 | Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature ConstancyabstractFor CNNs based stereo matching methods, cost volumes play an important role in achieving good matching accuracy. In this paper, we present an end-to-end trainable convolution neural network to fully use cost volumes for stereo matching. Our network consists of three sub-modules, i.e., shared feature extraction, initial disparity estimation, and disparity refinement. Cost volumes are calculated at multiple levels using the shared features, and are used in both initial disparity estimation and disparity refinement sub-modules. To improve the efficiency of disparity refinement, multi-scale feature constancy is introduced to measure the correctness of the initial disparity in feature space. These sub-modules of our network are tightly-coupled, making it compact and easy to train. Moreover, we investigate the problem of developing a robust model to perform well across multiple datasets with different characteristics. We achieve this by introducing a two-stage finetuning scheme to gently transfer the model to target datasets. Specifically, in the first stage, the model is finetuned using both a large synthetic dataset and the target datasets with a relatively large learning rate, while in the second stage the model is trained using only the target datasets with a small learning rate. The proposed method is tested on several benchmarks including the Middlebury 2014, KITTI 2015, ETH3D 2017, and SceneFlow datasets. Experimental results show that our method achieves the state-of-the-art performance on all the datasets. The proposed method also won the 1st prize on the Stereo task of Robust Vision Challenge 2018. Zhengfa Liang, Yulan Guo, Yiliu Feng, Wei Chen 0009, Linbo Qiao, Li Zhou 0009, Hengzhu Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Attention Unet++: A Nested Attention-Aware U-Net for Liver CT Image SegmentationabstractLiver cancer is one of the cancers with the highest mortality. In order to help doctors diagnose and treat liver lesion, an automatic liver segmentation model is urgently needed due to manually segmentation is time-consuming and error-prone. In this paper, we propose a nested attention-aware segmentation network, named Attention UNet++. Our proposed method has a deep supervised encoder-decoder architecture and a redesigned dense skip connection. Attention UNet++ introduces attention mechanism between nested convolutional blocks so that the features extracted at different levels can be merged with a task-related selection. Besides, due to the introduction of deep supervision, the prediction speed of the pruned network is accelerated at the cost of modest performance degradation. We evaluated proposed model on MICCAI 2017 Liver Tumor Segmentation (LiTS) Challenge Dataset. Attention UNet++ achieved very competitive performance for liver segmentation. Chen Li 0034, Yusong Tan, Wei Chen 0009, Xin Luo 0009, Yuanming Gao, Xiaogang Jia, Zhiying Wang 0003 |
ICIP | 3 |
| 2020 | CenterRepp: Predict Central Representative Point Set's Distribution For DetectionabstractObject detection has long been an important issue in the discipline of scene understanding. Existing researches mainly focus on the object itself, ignoring its surrounding environment. In fact, the surrounding environment provides abundant information to help detectors classify and locate objects. This paper proposes CRPDet, viz. CenterRepp Detector, a framework for object detection. The main function of CRPDet is accomplished by the CenterRepp module, which takes into account the surrounding environment by predicting the distribution of the central representative points. CenterRepp converts labeled object frames into the mean and standard variance of the sampling points' distribution. This helps increase the receptive field of objects, breaking the limitation of object frames. CenterRepp defines a position-fixed center point with significant weights, avoiding to sample all points in the surroundings. Experiments on the COCO test-dev detection benchmark demonstrates that our proposed CRPDet has comparable performance with state-of-the-art detectors, achieving 39.4 mAP with 51 FPS tested under single size input. Yu-Lin He, Limeng Zhang, Wei Chen 0009, Xin Luo 0009, Xiaogang Jia, Chen Li 0034 |
ICPR | 3 |
| 2020 | ANU-Net: Attention-based nested U-Net to exploit full resolution features for medical image segmentation
Chen Li 0034, Yusong Tan, Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Yuanming Gao |
Comput. Graph. | 3 |
| 2018 | Learning for Disparity Estimation Through Feature ConstancyabstractStereo matching algorithms usually consist of four steps, including matching cost calculation, matching cost aggregation, disparity calculation, and disparity refinement. Existing CNN-based methods only adopt CNN to solve parts of the four steps, or use different networks to deal with different steps, making them difficult to obtain the overall optimal solution. In this paper, we propose a network architecture to incorporate all steps of stereo matching. The network consists of three parts. The first part calculates the multi-scale shared features. The second part performs matching cost calculation, matching cost aggregation and disparity calculation to estimate the initial disparity using shared features. The initial disparity and the shared features are used to calculate the feature constancy that measures correctness of the correspondence between two input images. The initial disparity and the feature constancy are then fed into a sub-network to refine the initial disparity. The proposed method has been evaluated on the Scene Flow and KITTI datasets. It achieves the state-of-the-art performance on the KITTI 2012 and KITTI 2015 benchmarks while maintaining a very fast running time. Source code is available at http://github.com/leonzfa/iResNet. Zhengfa Liang, Yiliu Feng, Yulan Guo, Hengzhu Liu, Wei Chen 0009, Linbo Qiao, Li Zhou 0009 |
CVPR | 5 |
| 2017 | Megalloc: Fast Distributed Memory Allocator for NVM-Based ClusterabstractAs the expected emerging Non-Volatile Memory (NVM) technologies, such as 3DXPoint, are in production, there has been a recent push in the big data processing community from storage-centric towards memory-centric. Generally, in large-scale systems, distributed memory management through traditional network with TCP/IP protocol exposes performance bottleneck. Briefly, CPU- centric network involves context switching, memory copy etc. Remote Direct Memory Access (RDMA) technology reveals the tremendous performance advantage over than TCP/IP: Allowing access to remote memory directly bypassing OS kernel. In this paper, we propose Megalloc, a distributed NVM allocator exposes NVMs as a shared address space of a cluster of machines based-on RDMA. Firstly, it makes memory allocation metadata accessed directly by each machine, allocating NVM in coarse-grained way; secondly, adopting fine-grained memory chunk for applications to read or store data; finally, it guarantees high distributed memory allocation performance. Songping Yu, Nong Xiao 0001, Mingzhu Deng, Yuxuan Xing, Fang Liu 0002, Wei Chen 0009 |
NAS | 6 |
| 2017 | A high performance reliable NoC router
Lu Wang 0019, Sheng Ma, Chen Li 0015, Wei Chen 0009, Zhiying Wang 0003 |
Integr. | 4 |
| 2017 | Redesign the Memory Allocator for Non-Volatile Main MemoryabstractThe non-volatile memory (NVM) has the merits of byte-addressability, fast speed, persistency and low power consumption, which make it attractive to be used as main memory. Commonly, user process dynamically acquires memory through memory allocators. However, traditional memory allocators designed with in-place data writes are not appropriate for the non-volatile main memory (NVRAM) due to the limited endurance. In this article, first, we quantitatively analyze the wear-oblivious of DRAM-oriented designed allocator—glibc malloc and the inefficiency of wear-conscious allocator NVMalloc. Then, we propose WAlloc, an efficient wear-aware manual memory allocator designed for NVRAM: (1) decouples metadata and data management; (2) distinguishes metadata with volatility; (3) redirects the data writes around to achieve wear-leveling; (4) redesigns an efficient and effective NVM copy mechanism, bypassing the CPU cache partially and prefetching data explicitly. Finally, experimental results show that the wear-leveling of WAlloc outperforms that of NVMalloc about 30% and 60% under random workloads and well-distributed workloads, respectively. Besides, WAlloc reduces the average data memory writes in 64 bytes block by 1.5 times comparing with glibc malloc. With the fulfillment of data persistency, cache bypassing NVM copy is better than cache line flushing NVM copy with performance improvement circa 14%. Songping Yu, Nong Xiao 0001, Mingzhu Deng, Fang Liu 0002, Wei Chen 0009 |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2016 | InnerCache: A Tactful Cache Mechanism for RDMA-Based Key-Value StoreabstractHigh-Performance network technology, Remote Direct Memory Access (RDMA), has revealed its tremendous advantage over traditional TCP/IP. With its ultra-low latency and high bandwidth, RDMA has been extensively adopted in distributed environment, especially for in-memory key-value stores. However, although RDMA does provide the ability to interact with remote user space memory directly, memory copy still exists between data memory area and communication memory area with two-sided communication semantics in in-memory key-value store. In addition, using high performance one-sided communication semantics will expose memory totally, hence an inadvertent corrupt data operation could crash system. In this paper, we propose a tactful cache mechanism for RDMA-based in-memory key-value store -- InnerCache. Our design concerns two dimensions with respect to improve the system performance with two-sided communication semantics and make system less vulnerable. It merges one-sided and two-sided communication model through making communication memory cacheable. Experimental results show that InnerCache can efficiently improve the performance of RDMA-based in-memory key-value store. Songping Yu, Rujie Yu, Nong Xiao 0001, Fang Liu 0002, Wei Chen 0009 |
ICWS | 6 |
| 2016 | Shielding STT-RAM Based Register Files on GPUs against Read DisturbanceabstractTo address the high energy consumption issue of SRAM on GPUs, emerging Spin-Transfer Torque (STT-RAM) memory technology has been intensively studied to build GPU register files for better energy-efficiency, thanks to its benefits of low leakage power, high density, and good scalability. However, STT-RAM suffers from the read disturbance issue, which stems from the fact that the voltage difference between read current and write current becomes smaller as technology scales. The read disturbance leads to high error rates for read operations, which cannot be effectively protected by the SEC-DED ECC on large-capacity register files of GPUs. Prior schemes (e.g., read-restore) to mitigate the read disturbance usually incur either non-trivial performance loss or excessive energy overhead, thus not applicable for the GPU register file design that aims to achieve both high performance and energy-efficiency. To combat the read disturbance, we propose a novel software-hardware co-designed solution (i.e., Red-Shield ), which consists of three optimizations to overcome the limitations of the existing solutions. First, we identify dead reads at compiling stage and augment instructions to avoid unnecessary restores. Second, we employ a small read buffer to accommodate register reads with high-access locality to further reduce restores. Third, we propose an adaptive restore mechanism to selectively pick the suitable restore scheme, according to the busy status of corresponding register banks. Experimental results show that our proposed design can effectively mitigate the performance loss and energy overhead caused by restore operations while still maintaining the reliability of reads. Xuhao Chen 0001, Nong Xiao 0001, Lei Wang 0011, Fang Liu 0002, Wei Chen 0009, Zhiguang Chen 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2015 | RAID-6Plus: A Fast and Reliable Coding Scheme Aided by Multi-failure Degradation
Mingzhu Deng, Nong Xiao 0001, Songping Yu, Wei Chen 0009, Zhiguang Chen 0001, Fang Liu 0002 |
APSCC | 5 |
| 2015 | WAlloc: An efficient wear-aware allocator for non-volatile main memoryabstractThe non-volatile memory (NVM) has the illustrious merits of byte-addressability, fast speed, persistency and low power consumption, which make it attractive to be used as main memory. Commonly, user process dynamically acquires memory through memory allocators. However, traditional memory allocators designed with in-place data writes are not appropriate for non-volatile main memory (NVRAM) due to the limited endurance. For instance, the number of write operations is merely 108 times per PCM cell. In this paper, we quantitatively analyze the wear-oblivious of DRAM-oriented designed allocator-glibc malloc and the inefficiency of wear-conscious allocator-NVMalloc. For example, the average imbalance factor (the maximum/the average) of memory allocation is about 7.5 and 3, respectively. Based on our observations, we propose WAlloc, an efficient wear-aware manual memory allocator designed for NVRAM, decouples metadata and data, uses Less Allocated First Out allocation policy and redirects the data writes. Experimental results show that the wear-leveling of WAlloc outperforms that of NVMalloc about 30% and 60% under random workloads and well-distributed workloads, respectively. In addition, considering the trade-off between space and wear-leveling, WAlloc reduces average data memory writes in 64 bytes block by average 1.5X comparing with malloc with almost 8% extra space overhead. Songping Yu, Nong Xiao 0001, Mingzhu Deng, Yuxuan Xing, Fang Liu 0002, Zhiping Cai, Wei Chen 0009 |
IPCCC | 7 |
| 2015 | HIFFS: A Hybrid Index for Flash File SystemabstractFlash memory, especially NAND flash memory, has become a popular alternative for the design of storage system. Index schemes of conventional file systems do not take the flash memory characteristics into account and will cause poor performance. The current flash file systems work well only in the case of small capacity. To address the problem, we propose a new hybrid indexing scheme both in directory structure and file data index for NAND flash file system in this paper, which is called HIFFS (a Hybrid Index for Flash File System). HIFFS contains two components: hash tree directory and adaptive file data index. The hash tree directory uses a hash-based index to get better update performance and search efficiency. The adaptive file data index uses different file index strategy according to file size. The experiment results show that HIFFS outperforms the state-of-the-art file systems both in throughput. Xiaoquan Wu, Nong Xiao 0001, Fang Liu 0002, Wei Chen 0009 |
NAS | 5 |
| 2014 | Improving Speculation Accuracy with Inter-thread Fetching Value Prediction
Li Shen 0007, Zhiying Wang 0003, Hui Guo 0004, Wei Chen 0009 |
ICA3PP (2) | 6 |
| 2014 | Binary compatibility for embedded systems using greedy subgraph mapping
Xuhao Chen 0001, Li Shen 0007, Zhiying Wang 0003, Wei Chen 0009 |
Sci. China Inf. Sci. | 5 |
| 2013 | HEUSPEC: A Software Speculation Parallel ModelabstractConventional software speculative parallel models are facing challenge due to the increasing number of the processor core and the diversification of the application. The performance of the guest program under the software speculative parallel execution model is closely related to the speculation accuracy, the control overhead and the rollback overhead of the model. In order to improve the speculative accuracy and the load balance, as well as improve the overhead of the conventional model, in this paper, we proposed a novel speculative parallel model named HEUSPEC. The HEUSPEC includes 2 key techniques, the heuristic value prediction(HVP) and the dynamic task granularity resizing(DTGR). We have implemented the runtime system of the model in ANSI C language. The experiment results show that when the speedup of the HEUSPEC model can reach 4.51 on the average (12% higher than conventional model) when speculative depth equals to 7. Besides, it shows good scalability and lower memory cost. Li Shen 0007, Zhiying Wang 0003, Hui Guo 0004, Wei Chen 0009 |
ICPP | 6 |
| 2012 | An approach to minimizing the interpretation overhead in Dynamic Binary Translation
Wei Chen 0009, Dan Chen 0001, Zhiying Wang 0003 |
J. Supercomput. | 1 |
| 2011 | A Novel Chaining Approach to Indirect Control Transfer Instructions
Wei Chen 0009, Zhiying Wang 0003, Qiang Dou, Yongwen Wang |
ARES | 1 |
| 2011 | Characterizing Fine-Grain Parallelism on Modern Multicore PlatformabstractSince chip multiprocessors have dominated the processor market, developing a parallel programming model with proper trade-off between productivity and efficiency become increasingly important. As a typical fine-grain parallelism model, Intel Threading Building Blocks (TBB) simplifies parallel programming by runtime schedule. Despite its simplicity, it costs non-trivial runtime overhead which may increase as the thread counts increase. In this work, we conduct an experiment on real commodity hardware to evaluate performance scalability of TBB using PARSEC benchmark suite. We first compare TBB with Pthreads to show that TBB applications can achieve comparable performance as Pthreads applications. To find the performance bottleneck of TBB applications, we measure the runtime overhead of TBB focused on 3 basic TBB runtime activities. The result provides valuable implications which can be used to develop scalable runtime libraries and architectural support for alleviating performance bottlenecks. Xuhao Chen 0001, Wei Chen 0009, Li Shen 0007, Zhiying Wang 0003 |
ICPADS | 2 |
| 2010 | DSS: Applying asynchronous techniques to architectures exploiting ILP at compile timeabstractEmbedded application environments require both high performance and low power. Architectures exploiting instruction-level parallelism (ILP) at compile time, such as very long instruction word (VLIW) and transport triggered architecture (TTA), may satisfy the requirements. They can be further enhanced by using asynchronous circuits to significantly reduce power consumption. As such, we are interested in asynchronous processors with architectures exploiting ILP at compile time. However, most of the current asynchronous processors are based on RISC-like architectures. When designing asynchronous VLIW or TTA processors, the distribution of control introduces some serious problems, and errors may occur because of the variable latencies of operations. This paper investigates the asynchronous processor with architecture exploiting ILP at compile time. In order to overcome these problems, we propose a data source selecting (DSS) scheme to guarantee instructions run correctly on asynchronous VLIW and TTA processors. Concretely, an asynchronous pipelined processor based on TTA is designed. The micro-architecture of the proposed asynchronous TTA processor is presented and an asynchronous processor named Tengyue is implemented using 180nm technology. The experimental results, for a range of benchmarks and working modes, show that the implemented asynchronous TTA processor with DSS scheme support runs correctly and power dissipation is reduced to about 43% to 65% of the equivalent synchronous processor. Zhiying Wang 0003, Hongguang Ren, Wei Chen 0009, Hongyi Lu |
ICCD | 5 |
| 2010 | Vapor: Virtual Machine Based Parallel Program Profiling FrameworkabstractIt is hard to execute parallel program efficiently on man-core platform because we could not divide program into appropriate granularity executed simultaneously. Based on virtual machine and binary translation technologies the article proposes the vapor profiling framework that uses SBIRP instruction in-place replacement method to collect program's run-time control flow and data flow information precisely. Moreover, it explains how to create control flow and data flow dependency graphs. Experiment results prove that vapor has better performance than traditional methods. Yusong Tan, Wei Chen 0009, Qingbo Wu 0003 |
ICPADS | 2 |
| 2010 | A Novel Chaining Approach for Direct Control Transfer InstructionsabstractSoftware-based code cache systems are the key element in the dynamic translation system or optimization system to store the translated or optimized code for reuse. Translated code is organized in terms of code blocks in the code cache which transfers execution to the next code block through a control transfer instruction. As the target address of the control transfer instruction is in the form of its source program counter, the code cache system has to check the address mapping table for the translated program counter of the target address before entering the required code block. This will cause the performance degradation. As the target address of the direct control transfer instruction is fixed during the execution of a program, its source target address can be replaced with the translated target address. A direct control transfer chaining approach which occupies specific software assists is proposed in this paper. Evaluation of DCTC is conducted on a code cache simulator. The experiment results show the dramatic performance improvement brought by DCTC. Weixia Xu 0001, Wei Chen 0009, Qiang Dou |
ICPADS | 2 |
| 2009 | A Light-weight Code Cache Design for Dynamic Binary TranslationabstractInterpretation and basic block translation (BBT) are two typical strategies for cold code emulation in a dynamic binary translation (DBT) system. More and more DBT systems employ BBT as the generated native code runs more efficient than the interpretation routines. We observe that BBT's high efficiency is based on those special hardware assists. With certain simple hardware techniques, interpretation could outperform BBT. In our pervious work, we proposed a hardware interpreted code cache (Pcache) mechanism to speedup interpretation by saving the decoded instruction information during interpretation. This light-weight code cache design could be extended to assist the hotspots translation, thus further reduce the DBT systems' overhead. We add the translation entry into the Pcache design thus saving most decoding operations during translation. We use eight SPEC 2000 integer benchmarks on our DBT simulator. Results show that the modified Pcache design causes a speedup of 1.94 according to the referenced DBT with basic interpretation and the interpretation based DBT system assisted by the modified Pcache performs more efficiently than the DBT system which employs BBT for the cold code. Wei Chen 0009, Li Shen 0007, Hongyi Lu, Zhiying Wang 0003, Nong Xiao 0001 |
ICPADS | 1 |
| 2009 | Using Pcache to Speedup Interpretation in Dynamic Binary TranslationabstractAbstract— Dynamic binary translation (DBT) converts codes written for a source instruction set architecture (ISA) into optimized code for a target ISA. DBT has emerged as an important tool with real world applications. Interpretation is always adopted to handle the non-hotspot code in a two-stage DBT system. An important consideration in such DBT systems is the interpretation overhead. We investigate that repeated redecoding operations are the bottleneck of interpretation overhead. We propose interpreted code cache (Pcache), a hardware assist to save the information of the decoded instruction for reuse. We analyze and model Pcache performance via simulation on a DBT system simulator. Results from SPEC2000 integer benchmarks show that Pcache could significantly reduce redecoding operations and the overhead of interpretation in a DBT system. The speedup of interpretation is up to 17.12 on average with assist of Pcache. We also analyze the extra overhead caused by Pcache, which is neglectable compared to the performance gains. Wei Chen 0009, Hongyi Lu, Li Shen 0007, Zhiying Wang 0003, Nong Xiao 0001 |
ISPA | 1 |
| 2008 | A New Approach to Single Event Effect Tolerance Based on Asynchronous Circuit Technique
Wei Chen 0009, Fang Liu 0002, Kui Dai, Zhiying Wang 0003 |
J. Electron. Test. | 2 |