Gang Dong

dblp:84/2469 · DBLP profile ↗
← Back
38ranked-venue papers
7as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 6 · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Robust Exemplar Prompt Learning via Bi-directional Visual-Semantic Alignment for Multi-Object Tracking
abstract
Recent multi-object tracking (MOT) approaches increasingly leverage pre-trained CLIP models to boost cross-domain generalization. A common strategy uses a predefined TrackBook—a closed-set of visual concepts—as textual prompts to guide learning of domain-invariant representations. However, these fixed prompts lack adaptive context, causing limited generalization. To address this limitation, this paper introduces a robust Exemplar Prompt Learning (EPL) framework via Bi-directional Visual-Semantic Alignment (BiVSA), termed EPL-MOT, which augments textual prompts with instance-aware contextual information derived during tracking. Specifically, an EPL module is designed to dynamically enrich textual prompts with contextual cues, enabling instance-specific adaptation without inducing category shift. Furthermore, a BiVSA module is proposed to deepen cross-modal interaction by incorporating bidirectional learnable prompts into both textual and visual branches. This facilitates progressive integration of global semantic features with local visual structures, resulting in a more effectively aligned visual-semantic space. Finally, to enhance robustness against distractors, a Category-guided Detection Query Generator (CDQG) is constructed, which incorporates base-class textual information to suppress irrelevant targets. Comprehensive evaluations on MOT17 and MOT20 demonstrate that the proposed EPL-MOT achieves competitive performance across both in-domain and cross-domain settings.
Lingyan Liang, Gang Dong, Dongchao Wen, Kaihua Zhang 0001
ICMR3
2026 FPGA-based heterogeneous computing framework for environment-level parallel decision-making in autonomous driving and robotic control
Hongbin Yang 0003, Yaqian Zhao, Ruyang Li, Gang Dong
Expert Syst. Appl.5
2025 Continuously Learning Video-level Object Tokens for Robust UAV tracking
abstract
Due to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the object template and the search frame. The drastic object appearance variations degrade the learned model, leading to drift issue. To this end, this paper presents a video-level UAV tracking framework that focuses on Continuously Learning (CL) effective and efficient spatio-temporal object tokens for robust tracking, dubbed as CLTrack. Specifically, the CLTrack first learns a series of spatio-temporal object tokens via a dynamic filtering module (DFM), which encodes more consensus object appearance information from each frame. Afterwards, a spatio-temporal enhancement module (STEM) is designed via cascading a temporal and a spatial attention to fully interact with the selected tokens with stable long-range spatio-temporal context information of the tracked object. Finally, to ensure the learned model encodes the rich context information without catastrophic forgetting, a video-level tracking loss is designed to supervise feature learning from the whole video frames. Extensive experiments on three UAV benchmarks including UAV123, DTB70 and VisDrone2018 demonstrate that the proposed CLTrack achieves state-of-the-art performance.
Shenglong Hu, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001
ICASSP3
2025 Easy-to-hard Instance-level Feature Fusion for Co-saliency Detection
abstract
Existing leading deep learning-based Co-saliency Detection (CoD) methods often learn the consensus features from the input image group without considering the complexity of each image. Despite the demonstrated success, the input images may contain hard samples with high complexity, e.g., those containing distractors that have similar appearance but different semantics to the co-salient objects. This is prone to mislead the learned model to treat these distractors as co-salient objects, leading to classification ambiguity. To address this issue, this paper presents an easy-to-hard instance-level feature Fusion framework for CoD, termed E2HCoD. The E2HCoD exploits the instance-level co-salient object consensus cues from the easy samples as reliable guidance to accurately fuse the co-salient object features in the hard samples. First, we design a Feature Filtering Module (FFM) that evaluates image complexity by integrating entropy, variance, texture, and edge density cues, allowing the model to select the easy samples with relatively easy backgrounds. Then, we develop an Easy-instance Embedding Branch (EEB), which accurately segments the co-salient object masks from the easy samples as the instance-level guidance to learn the accurate co-salient object consensus cues. Then, with the consensus knowledge from the easy samples as guidance, we construct an Easy-instance guided Fusion Branch (EFB), which fully interacts with the consensus features from the hard samples via a cross-attention mechanism, yielding the refined features that highlight the co-salient objects while suppressing the distractors. Finally, the refined features are fed into the decoder, generating a high-quality CoD prediction. Extensive experiments demonstrate that the proposed E2HCoD achieves state-of-the-art performance on CoSal2015, CoCA, and CoSOD3k.
Chuang Ding, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001
ICASSP3
2025 Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change Detection
abstract
Existing leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS images, this paradigm is prone to be misled by irrelevant semantic category changes, leading to a noisy CD mask prediction. To address this issue, this paper presents a Spatio-Semantic Prompt (SSP) guided adaptive Segment Anything Model (SAM) for RSCD, dubbed as SSP-SAM. The SSP-SAM introduces sparse textual and dense mask prompts into SAM to encode the task-specific semantic knowledge for RSCD. Specifically, we first encode the powerful textual semantic knowledge using Contrastive Language-Image Pre-training (CLIP) to determine the desired change semantic category. Then, we design a spatial dense prompt module that yields an attention map as prompt features to further refine the desired changed regions. Subsequently, we fine-tune the SAM through an adaptor to integrate the spatial-semantic prompt cues, yielding a coarse CD mask prediction. Finally, guided by the coarse CD mask, a multi-scale mask attention mechanism is adopted to learn the refined semantic representations of the changed targets, predicting the accurate CD mask. Extensive experiments on a variety of benchmark datasets demonstrate that the proposed SSP-SAM achieves state-of-the-art performance.
Shenglong Hu, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001
ICASSP3
2025 Tripartite interaction representation learning for multi-modal sentiment analysis
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Wenfeng Yin
Expert Syst. Appl.2
2025 Transmission Channel and Matching Network Design Strategies in Chiplet-Based 2.5-D ICs for RF Applications
abstract
In this paper, a neural network-based approach for the co-design of transmission channels and matching networks in chiplet-based 2.5-D integrated circuits (ICs) for radio frequency (RF) applications is presented. This work proposes a dual-neural network framework that synergizes cascaded S-parameter modeling with data augmentation techniques. First, a cascaded method is developed to accurately model the 2.5-D signal transmission channel, incorporating through-silicon vias (TSVs), redistribution layers (RDLs), and matching networks, validated against commercial circuit tools. A convolutional neural network (CNN) is then trained to inversely map S-parameter requirements to optimal structural and matching network parameters. Concurrently, a denoising diffusion probabilistic model (DDPM) generates augmented S-parameter datasets to expand the design space while reducing simulation overhead. The framework enables automated compliance with critical constraints while optimizing secondary objectives such as matching network area minimization and frequency response enhancement. This work provides a robust artificial intelligence (AI)-driven methodology for high-efficiency RF interconnects in heterogeneous integration platforms.
Changle Zhi, Gang Dong, Deguang Yang, Junpeng Yao, Daihang Liu, Zhangming Zhu
IEEE Internet Things J.2
2025 Miniaturized diplexer with wide-stopband based on half-mode substrate integrated waveguide
abstract
基于半模基片集成波导(HMSIW),提出一种宽阻带小型化双工器。该双工器将双模谐振器(DMR)与单模谐振器(SMR)相结合。采用HMSIW技术突破了SMR在小型化方面的限制,同时有效解决了SMR中常见的TE202模式对宽阻带性能的限制。采用TE101/TE301 DMR设计并制造了一个二阶原型,原型的中心频率分别为10.34 GHz和13.90 GHz。测量结果显示,该双工器的阻带范围为2.16f1(f1是通道1的中心频率),具有优于20 dB的带外抑制水平。与此同时,该双工器的尺寸被显著缩小至1.363λg2(λg是f1时介质基底中的导波波长)。
Gang Dong, Xinqing Lei, Zhangming Zhu
Frontiers Inf. Technol. Electron. Eng.2
2025 FrigateBird: Decoupling Metadata/Data Services for Continuously Fast Object Storage
abstract
Currently, ride-hailing service has tens of millions of registered drivers and hundreds of millions of registered passengers, serving tens of millions of rides per day. Different from social media applications like Facebook and LinkedIn, ride hailing needs to read/write/query a large number of small objects (like photos and audio/video pieces)always fast, so that it can support critical online computations such as face comparison and sentiment analysis on audio/video records. This is of particular importance for ride-hailing service to recognize and avoid potential dangers. Existing object stores (like Haystack and Tectonic) usually store object data in files and place object metadata in a separate key-value store (like RocksDB), which is unsuitable for ride-hailing service mainly because the objects' file-related information is placed together with the object data. This severely affects the I/O performance of object storage: first, for crash consistency, the writes of object data and object metadata must be conducted inseparatephases of one transaction, which significantly increases I/O latency; second, the space of deleted objects needs to be reclaimed viacompaction, which could sharply lower I/O performance when the system is busy in serving normal read/write requests. This paper describes FrigateBird, a continuously fast object store for ride-hailing service. FrigateBird differs from existing object stores in three aspects. First, we present a metadata/data decoupled service architecture for object storage, where therichmetadata service realizes efficient queries and updates of object metadata, and therawdata service purely performs disk I/O to read/write object data from/to raw disks. Second, we propose a rich metadata structure (calledRichmeta) taking the write operation logs as part of object metadata, which allows FrigateBird tosimultaneouslywrite the object data to raw disks (without filesystem overhead) and write the metadata to a key-value store, guaranteeing crash consistency by checking whether the transaction is completed and rolling back if not. Third, we design a compaction-free deletion mechanism which can efficiently delete an object by only updating the metadata without involving the data service, so that FrigateBird can efficiently support ride-hailing service's frequent delete operations while avoiding data-migration-caused performance hiccup. Evaluation shows that FrigateBird outperforms the state-of-the-art object stores by up to$3.36\times$and$27.9\times$in the mean I/O latency for normal and in-compaction scenarios, respectively.
Yiming Zhang 0003, Ke-Kun Hu, Gang Dong, RenGang Li
IEEE Trans. Serv. Comput.4
2025 Electrical and Thermal Characteristics Optimization in Interposer-Based 2.5-D Integrated Circuits
abstract
In this work, a comprehensive analysis and optimization method of electrical and thermal characteristics in 2.5-D integrated circuits (ICs) is performed, including rapid heat distribution modeling, integrated voltage regulator (IVR) chip modeling, and power delivery network (PDN) modeling. Based on the proposed method, chiplet placement, decoupling placement, and IVR parameter settings that compromise the total PDN impedance, IVR impedance, and thermal distribution characteristics can be obtained. First, a rapid thermal analysis method for multiple heat sources is proposed by integrating the equivalent thermal resistance method and commercial tools. The thermal method significantly improves the computational efficiency and reduces the memory usage. Then, we analyze the electrical characteristics of a typical low dropout (LDO) and model the complete 2.5-D PDN, including interposers, chiplets, IVRs, through-silicon vias (TSVs), bumps, decoupling capacitors, and other components. The electrical and thermal problems in the 2.5-D system are formulated and a Metropolis rule-based algorithm is used to derive optimal solutions. Finally, the optimal placement schemes and parameter settings are iterated under different constraints. This method allows for the adjustment of target impedance, noise current, thermal limit, and other constraints based on varying practical situations. In the time-domain analysis, it can be found that the capacitance value is reduced while maintaining the power supply performance. With high accuracy in thermal and electrical modeling, this work provides an in-depth reference for the co-design of chiplet-based 2.5-D ICs.
Changle Zhi, Gang Dong, Deguang Yang, Daihang Liu, Yinghao Feng, Yang Wang 0104, Zhangming Zhu
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change Detection
abstract
Existing deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and slow convergence of the model. To this end, we propose a novel segment anything model (SAM) guided framework, termed as SAM-CD, which mines the rich semantic knowledge from the SAM for RSCD. Specifically, we first employ a transformer encoder to extract multi-scale global features from the bi-temporal images. Meanwhile, we obtain semantic prior masks from the bi-temporal images by providing the SAM with category-relevant text prompts. Then, using the semantic prior masks as constraints, we design a masked attention module (MAM) that generates local features related to the interested categories. Finally, the local and global features are fused and fed into a multi-layer perception (MLP) decoder to obtain the change map. The whole network is trained in an end-to-end manner that can readily encode the rich semantic knowledge of the changed targets to predict an accurate change map. Extensive experiments demonstrate that the proposed SAM-CD achieves state-of-the-art performance on a variety of benchmark datasets.
Zixuan Sun, Huihui Song 0003, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao
ICASSP4
2024 Glance, Focus and Refinement Network for Remote Sensing Change Detection
abstract
Existing change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively capture the change targets due to the imbalance ratio between the change regions and the whole scene. To this end, this paper presents a glance, focus, and refinement network (GFRNet), which formulates CD as a continuous, step-by-step focusing process to mimic the human visual system. Specifically, the GFRNet first employs a transformer encoder to extract the global features from the bi-temporal images, where each feature takes a glance at the whole scene. Then, the GFRNet gradually pays attention to a cascade of salient regions, and ultimately progressively refines its focus on the desired areas of change. Comprehensive evaluations on two extensively utilized benchmark datasets, including LEVIR-CD and WHU-CD, demonstrate the superiority of our GFR-Net to a variety of state-of-the-art methods.
Zixuan Sun, Yuhui Zheng, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao
ICASSP5
2024 Group-wise co-salient object detection via multi-view self-labeling novel class discovery
Gang Dong, Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001
Frontiers Comput. Sci.2
2024 Easy Pruning via Coresets and Structural Re-Parameterization
abstract
Inference time pruning is characteristic in high construction efficiency, since it dramatically reduces the dependency on finetuning to recover precision. It is adequate to reconstruct compressed convolution kernels by optimizing the loss of feature map reconstruction. However, the accuracy decline of compressed network increases as the loss of feature map reconstruction accumulates layer by layer. To enhance layerwise convolution kernel reconstruction, this paper proposes a hybrid method via combining coresets theory and structural re-parameterization, enabling shallow transfer learning (STL) during inference time pruning. Firstly, our method achieves STL by implementing structural re-parameterization in the process of convolution kernel reconstruction, to adapt to effects of one layer's reconstruction loss on the next layers' inputs. Secondly, a channel-wise scaling process is designed on the basis of coresets theory, to enhance approximation in the mapping from drifted inputs to original feature maps. Selectively, a maximum mean discrepancy based decision-making process is built for switching in two patterns of our method. Tests are executed on image classification and arrhythmia detection. As observed on ImageNet datasets, coresets theory based scaling is more effective at filter level for DenseNet and MobileNet-v2 and resultful at unit kernel level for ResNet and SqueezeNet.
Wenfeng Yin, Gang Dong, Dianzheng An, Yaqian Zhao, Binqiang Wang
IEEE Signal Process. Lett.2
2024 Multiobjective Optimization for PSIJ Mitigation and Impedance Improvement Based on PCPS/DR-NSDE in Chiplet-Based 2.5-D Systems
abstract
The utilization of modular chiplets in interposer-based 2.5-D heterogeneous systems simplifies fabrication and design, however, it also introduces significant noise challenges. This paper presents a collaborative jitter-aware optimization in 2.5-D integrated circuits (ICs), incorporating power supply induced jitter (PSIJ), system impedance, target impedance, and decoupling capacitors, based on the hybrid pre-computation and pre-storage/duplicate removal-non-dominated sorting differential evolution (PCPS/DR-NSDE) algorithm. An automatic channel model algorithm and a uniform decoupling capacitor placement strategy are proposed to improve the design efficiency. Then, the system transfer impedance, simultaneous switch current, sensitivity function, and amplification factor are individually modeled, leading to the assembly and verification of the final PSIJ in the 2.5-D system. A PCPS strategy is proposed to handle high-time-consuming modules in the objective function and a DR operation is added to improve algorithm performance. The proposed PCPS/DR-NSDE is faster than traditional algorithms and has optimal hypervolume and coverage-metric (C-metric) indicators. The procedures for further obtaining desired solutions in the Pareto front are discussed. The impact of practical constraints and target impedance is also analyzed. This work provides a collaborative optimization and analysis of jitter, noise, and impedance in 2.5-D systems.
Changle Zhi, Gang Dong, Deguang Yang, Daihang Liu, Yinghao Feng, Yang Wang 0104, Zhangming Zhu, Yintang Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Coresets based asynchronous network slimming
abstract
Abstract Pruning is effective to reduce neural networks’ parameters and accelerate inferences, facilitating deep learning in resource-limited scenarios. This paper proposes an asynchronous pruning method for multi-branch networks on the basis of our previous work on channel coresets constructions, to achieve module-level pruning. Firstly, this paper accelerates coreset based pruning by batch sampling with a sampling probability decided on our-designed importance function. Secondly, this paper gives asynchronous pruning solutions with an in-place distillation of feature maps for deployment on multi-branch networks such as ResNet and SqueezeNet. Thirdly, this paper provides an extension to neuron pruning by grouping weights as channels. During tests on sensitivity of different layers to channel pruning, our method outperforms comparison schemes on object detection networks, indicating advantages of data-independent channel selections in maintaining precision. As shown in tests of asynchronous pruning solutions on multi-branch classification networks, our method further decreases FLOPs with a small accuracy decline on ResNet and acquires a small accuracy increment on SqueezeNet. In tests on neuron pruning, our method achieves an accuracy comparable to existing coreset based pruning methods by two solutions of precision recovery.
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li
Appl. Intell.2
2023 Correction to: Coresets based asynchronous network slimming
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li
Appl. Intell.2
2023 Stacked arrangement of substrate integrated waveguide cavity-backed semicircle patches for wideband circular polarization with filtering effect
abstract
文章提出一种应用于 X 频段和 Ku 频段卫星无线通信的带有滤波效应的新型宽带圆极化天线。该结构包含一个驱动层(同时也是滤波层)以及一个堆叠层(同时也是圆极化层)。带通滤波响应中的两个辐射零点,是衬底集成波导(SIW)腔体支持的开口与嵌入式驱动贴片的综合效果。引入倒角贴片作为堆叠元件,具有同时实现圆极化和拓宽工作带宽的能力。使用多层印制电路板(PCB)工艺制作了一个尺寸为 0.8λ0×0.71λ0×0.16λ0 的紧凑原型进行演示。实验结果与仿真结果吻合良好,测量的 −10-dB 阻抗带宽和 3-dB 轴比带宽分别为10.83% 和 15.54%。此外,还获得了 8.9 dBic 的左旋圆极化峰值增益,大于 7 dBic 的带内平均左旋圆极化增益,以及良好的频率选择性。
Yitong Yao, Gang Dong, Zhangming Zhu, Yintang Yang
Frontiers Inf. Technol. Electron. Eng.2
2023 Hierarchically stacked graph convolution for emotion recognition in conversation
abstract
Accurate emotion recognition can drive the robot to understand human affection intentions precisely and deliver the emotional response when communicating with a person. Recently, graph structure has been applied to explicitly capture the self and inter-dependencies of speakers in the conversation. However, the performance of the method is limited by inadequate discriminative information extraction based on naive graph convolution. In this paper, we propose a novel Hierarchically Stacked Graph Convolution Framework (HSGCF), which leverages hierarchical structure to extract emotional discriminative features. The proposed HSGCF uses five graph convolution layers connected hierarchically to establish a more discriminative emotional feature extractor. More importantly, to mitigate the over-smooth problem caused by deeper networks, Transformer structures with residual connection are introduced into HSGCF. Experimental results on the IEMOCAP benchmark dataset indicate the proposed framework achieves a 4.12% improvement in accuracy and a 4.80% improvement in F1 score compared with the baseline method.
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Ke-Kun Hu, Dongdong Jiang
Knowl. Based Syst.2
2023 FY-3E Wind Scatterometer Prelaunch and Commissioning Performance Verification
abstract
The spaceborne microwave scatterometer (SCAT) is a radar for quantitatively measuring the backscatter coefficient of the Earth’s surface. Its main application is the measurement of wind speed and direction near the sea surface with a wide swath and high precision. So far, many SCATs have been launched in orbit, including SeaWinds by the United States, the ASCAT series by Europe, and the HY-2 series scatterometers by China, which contribute to improved weather forecasts and typhoon positioning services. The Fengyun-3E satellite wind scatterometer or WindRad, launched by China in July 2021, is the first Ku and C dual-band SCAT with the highest spatial resolution. In this study, the rotation characteristics and calibration parameters of WindRad in the prelaunch stage and after its early test and verification in orbit were investigated. The prelaunch test results demonstrated the excellent performance of WindRad, which can fully meet its quantitative application requirements. In-orbit early tests demonstrated the stable operation of WindRad, with signal characteristics consistent with the design specifications. The Ku and C dual-band earth surface backscatter coefficient projection and the sea surface wind field inversion were also investigated. The preliminary results confirmed the ability of WindRad to provide a high-quality ocean wind measurement and the improved capability for different applications, including sea ice detection and soil moisture measurement.
Axin Jin, Anzhong Jin, Shiyu Xue, Pei Yao, Gang Dong, Haoqiang Shi, Ailing Lv, Jian Shang
IEEE Trans. Geosci. Remote. Sens.6
2023 Miniaturization Strategy for Directional Couplers Based on Through-Silicon Via Insertion and Neuro-Transfer Function Modeling Method
abstract
Enhancing the integration of directional couplers is a crucial challenge in the design of wireless communication circuits and systems. This article proposes a design strategy based on through-silicon via (TSV) insertion and neuro-transfer function (Neuro-TF) modeling to achieve miniaturization of coupled-line couplers. By embedding coupled TSV pairs inside the substrate, a planar coupler can be transformed into a 3-D structure with a smaller footprint. Furthermore, a design flow is proposed to achieve the initial parameters of TSV insertion for various application scenarios. Two insertion schemes with regular and coaxial TSVs are applied to handle different requirements and process constraints. The flow also estimates the number of required TSV pairs based on the area reduction target. To reduce dependence on electromagnetic (EM) simulation and make the strategy more widely applicable to various coupler types, Neuro-TF modeling is utilized for optimization after the initial design. Subsequently, the compatibility between the TSV insertion flow and the Neuro-TF method is analyzed to develop a comprehensive miniaturization strategy. The approach has been validated by two design cases through EM simulations. The results indicate that the proposed strategy enables rapid design and achieves area reduction targets effectively.
Gang Dong, Changle Zhi, Yang Wang 0104, Zhangming Zhu, Yintang Yang
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Learning from Fourier: Leveraging Frequency Transformation for Emotion Recognition
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li
ICONIP (2)2
2022 Non-Uniform Attention Network for Multi-modal Sentiment Analysis
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Yinyin Chao
MMM (1)2
2022 Trade-Off-Oriented Impedance Optimization of Chiplet-Based 2.5-D Integrated Circuits With a Hybrid MDP Algorithm for Noise Elimination
abstract
Interposer and chiplet-based 2.5-D integrated circuit (IC) designs have become a new trend for block-level heterogeneous integration. In this paper, a new hybrid metaheuristic algorithm named Metropolis-based differential particle swarm optimization (MDP) is designed to jointly optimize the multiconstraints and impedance-based hybrid objective function of chiplet-based 2.5-D IC including interposers, chiplets, through-silicon via (TSV) arrays, bumps, and metal-insulator-metal (MIM) capacitors for simultaneous switch noise (SSN) reduction. Combined with the cascaded PDN assembly method, constraints on routing, delay and proximity distance between the entire system and an impedance-oriented function with multiple critical factors, a hybrid objective function with respect to the 2.5-D PDN is obtained. Integrating the advantages of multiple algorithms, a better hybrid MDP algorithm is designed to optimize the proposed key function. This method adopts the Metropolis rule to avoid the waste of the update mechanism for out-of-boundary particles. The placement, orientation of the chiplets, the on-interposer decoupling capacitor and the constraints of the 2.5-D system are co-optimized to find the optimal solution to eliminate the SSN. The overdesign of the system, different target impedance, different objective-oriented circuit optimization schemes and trade-offs in different constraints are also discussed carefully in this paper for 2.5-D ICs.
Changle Zhi, Gang Dong, Yang Wang 0104, Zhangming Zhu, Yintang Yang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 3-D Compact Marchand Balun Design Based on Through-Silicon via Technology for Monolithic and 3-D Integration
abstract
An original concept of 3-D through-silicon via (TSV)-based Marchand baluns and its design methodology is proposed for on- chip and 3-D integration. By utilizing a 3-D packaging process that includes a TSV and redistribution layer (RDL), a meander coupling path is established and embedded in the vertical direction of the substrate for balun design. The 3-D structure can effectively reduce the on- chip area while maintaining good balance characteristics. Furthermore, the structure can be flexibly integrated with 3-D integrated circuits (3-D ICs) to realize signal conversion between different stacking tiers. An equivalent circuit model based on the TSV-to-TSV coupling channel and coupled transmission line has been established for initial estimation before electromagnetic (EM) optimization. To shorten the design cycle, a specific flow is proposed to meet the structural particularity and different application scenarios. To verify the design method and flow, a design case is analyzed and EM simulated with the antenna feeding network as the target application. The EM simulation results show that the design can work at 43–82 GHz with an amplitude imbalance of less than 0.3 dB and a phase imbalance of less than 1.3°, which meets the requirements of balanced feeding. It only costs a$0.084\,\,\lambda _{\text {g}}\,{\times }\,0.009\,\,\lambda _{\text {g}}$footprint, which is far lower than that of the conventional planar types.
Gang Dong, Yang Wang 0104, Zhangming Zhu, Yintang Yang
IEEE Trans. Very Large Scale Integr. Syst.2
2021 Coresets Application in Channel Pruning for Fast Neural Network Slimming
abstract
Pruning reduces neural networks' parameters and accelerates inferences, enabling deep learning in resource-limited scenarios. Existing saliency-based pruning methods apply characteristics of feature maps or weights to judge the importance of neurons or structures, where weights' characteristics based methods are data-independent and robust for future input data. This paper proposes a coreset based pruning method for the data-independent structured compression, aiming to improve the construction efficiency of pruning. The first step of our method is to prune channels, according to the channel coreset merged from multi-rounds coresets constructions. Our method adjusts the importance function utilized in the random probability sampling during coresets construction procedures to achieve data-independent channel selections. The second step is recovering the precision of compressed networks through solving the compressed weights reconstruction by linear least squares. Our method is also generalized to implementations on multi-branch networks such as SqueezeNet and MobileNet-v2. In tests on classification networks like ResNet, it is observed that our method performs fast and achieves an accuracy decline as small as 0.99% when multiple layers are pruned without finetuning. As shown in evaluations on object detection networks, our method acquires the least decline in mAP indicator compared to comparison schemes, due to the advantage of data-independent channel selections of our method in preserving precision.
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li
IJCNN2
2020 MEgATrack: monochrome egocentric articulated hand-tracking for virtual reality
abstract
We present a system for real-time hand-tracking to drive virtual and augmented reality (VR/AR) experiences. Using four fisheye monochrome cameras, our system generates accurate and low-jitter 3D hand motion across a large working volume for a diverse set of users. We achieve this by proposing neural network architectures for detecting hands and estimating hand keypoint locations. Our hand detection network robustly handles a variety of real world environments. The keypoint estimation network leverages tracking history to produce spatially and temporally consistent poses. We design scalable, semi-automated mechanisms to collect a large and diverse set of ground truth data using a combination of manual annotation and automated tracking. Additionally, we introduce a detection-by-tracking method that increases smoothness while reducing the computational cost; the optimized system runs at 60Hz on PC and 30Hz on a mobile processor. Together, these contributions yield a practical system for capturing a user's hands and is the default feature on the Oculus Quest VR headset powering input and social presence.
Shangchen Han, Randi Cabezas, Christopher D. Twigg, Peizhao Zhang, Jeff Petkau, Tsz-Ho Yu, Chun-Jung Tai, Muzaffer Akbay, Asaf Nitzan, Gang Dong, Yuting Ye, Lingling Tao, Chengde Wan, Robert Wang 0002
ACM Trans. Graph.12
2019 Thermal-Aware Modeling and Analysis for a Power Distribution Network Including Through-Silicon-Vias in 3-D ICs
abstract
The high-power dissipation and the low-heat conduction between stacked chips in 3-D structures lead to a high temperature, which has become a critical design constraint for high-performance 3-D integrated circuits (ICs). Naturally, a power distribution network (PDN) in 3-D ICs consists of good heat conductor, but many existing methods do not fully show its thermal removal capability, and the thermal models of a 3-D PDN are imperfect. In this paper, compact and accurate physically based models to compute effective thermal conductivities of a 3-D PDN are proposed. By combining the equivalent thermal conductivity and the proposed complete thermal models for a 3-D PDN, a fast temperature analysis procedure can be expediently applied using the finite volume method. Finite element method (FEM) tools are used to verify the accuracy of the proposed models, and the errors in the temperature responses of our proposed method are within 3% compared with the results from the FEM. In addition, our method also improves the computational efficiency.
Weijun Zhu, Gang Dong, Yintang Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Antenna-in-package system integrated with meander line antenna based on LTCC technology
abstract
We present an antenna-in-package system integrated with a meander line antenna based on low temperature co-fired ceramic (LTCC) technology. The proposed system employs a meander line patch antenna, a packaging layer, and a laminated multi-chip module (MCM) for integration of integrated circuit (IC) bare chips. A microstrip feed line is used to reduce the interaction between patch and package. To decrease electromagnetic coupling, a via hole structure is designed and analyzed. The meander line antenna achieved a bandwidth of 220 MHz with the center frequency at 2.4 GHz, a maximum gain of 2.2 dB, and a radiation efficiency about 90% over its operational frequency. The whole system, with a small size of 20.2 mm×6.1 mm×2.6 mm, can be easily realized by a standard LTCC process. This antenna-in-package system integrated with a meander line antenna was fabricated and the experimental results agreed with simulations well.
Gang Dong, Zhao-yao Wu, Yintang Yang
Frontiers Inf. Technol. Electron. Eng.1
2015 A Doubly Degenerate Diffusion Model Based on the Gray Level Indicator for Multiplicative Noise Removal
abstract
Multiplicative noise removal is a challenging task in image processing. Inspired by the impressive performance of nonlinear diffusion models in additive noise removal, we address this problem in the view of nonlinear diffusion equation theories rather than the traditional variation methods. We develop a nonlinear diffusion filter denoising framework, which considers not only the information of the gradient of the image, but also the information of gray levels of the image. Furthermore, under this framework, we propose a doubly degenerate diffusion model for multiplicative noise removal, which is analyzed with respect to some of its properties and behavior in denoising process. In numerical aspects, we present an efficient scheme which uses a stabilization by fast explicit diffusion for the implementation of the multiplicative noise removal model. Finally, the experimental results illustrate effectiveness and efficiency of the proposed model.
Zhichang Guo, Gang Dong, Jiebao Sun, Dazhi Zhang, Boying Wu
IEEE Trans. Image Process.3
2008 A Robust Method for Edge-Preserving Image Smoothing
Gang Dong, Kannappan Palaniappan
ACIVS1
2007 On the Convergence of Bilateral Filter for Edge-Preserving Image Smoothing
abstract
The bilateral filter represents a wide group of nonlinear filters for edge-preserving image smoothing. In this work, we study the convergence properties of the bilateral filter algorithm. The understanding is established that the bilateral filter is an optimization procedure. We demonstrate that the bilateral filter is equivalent to minimizing a robust cost criterion using iterative reweighting, which is a good approximation to the very fast but unstable Newton's method. Further, the results of the analysis allow us to derive an improved hybrid smoothing scheme with concerns of computational efficiency and edge preservation.
Gang Dong, Scott T. Acton
IEEE Signal Process. Lett.1
2006 Fingerprint Image Enhancement by Diffusion Processes
abstract
Fingerprint enhancement is one of the most important steps in automatic fingerprint identification system. In this paper, based on nonlinear diffusion processes, we present an enhancement method which is able to adaptively enhance the ridge and valley structures in fingerprint images using the local ridge orientation information. Instead of using conventional gradient-based orientation estimation method, an anisotropic diffusion scheme is used to improve the estimation accuracy. Experimental results show a significant performance improvement of the fingerprint identification system by incorporating the proposed enhancement method.
Huiqing Chen, Gang Dong
ICIP2
2006 Motion Flow Estimation from Image Sequences with Applications to Biological Growth and Motility
abstract
In this paper, a new method for motion flow estimation that considers errors in all the derivative measurements is presented. Based on the total least squares (TLS) model, we accurately estimate the motion flow in the general noise case by combining noise model (in form of covariance matrix) with a parametric motion model. The proposed algorithm is tested on two different types of biological motion, a growing plant root and a gastrulating embryo, with sequences obtained microscopically. The local, instantaneous velocity field estimated by the algorithm reveals the behavior of the underlying cellular elements.
Gang Dong, Tobias I. Baskin, Kannappan Palaniappan
ICIP1
2005 Tracking multiple cells by correspondence resolution in a sequential Bayesian framework
abstract
We propose a multi-target tracking (MTT) algorithm in a sequential Bayesian framework that computes cell velocities from video microscopy. Unlike the traditional tracking methods, our formulation does not involve the estimation of target states; instead, we estimate one-to-one target correspondences by way of a sequential Markov chain Monte Carlo (MCMC) algorithm. The proposed probabilistic framework also automatically accounts for a variable number of targets. We have tested the proposed tracking algorithm on two different in vitro and one in vivo microscopy experiments. The three experiments show that the method holds promise in terms of low false positive and false negative rates as well as low rates of correspondence error.
Nilanjan Ray, Gang Dong, Scott T. Acton
ICIP (1)2
2005 Intravital leukocyte detection using the gradient inverse coefficient of variation
abstract
The problem of identifying and counting rolling leukocytes within intravital microscopy is of both theoretical and practical interest. Currently, methods exist for tracking rolling leukocytes in vivo, but these methods rely on manual detection of the cells. In this paper we propose a technique for accurately detecting rolling leukocytes based on Bayesian classification. The classification depends on a feature score, the gradient inverse coefficient of variation (GICOV), which serves to discriminate rolling leukocytes from a cluttered environment. The leukocyte detection process consists of three sequential steps: the first step utilizes an ellipse matching algorithm to coarsely identify the leukocytes by finding the ellipses with a locally maximal GICOV. In the second step, starting from each of the ellipses found in the first step, a B-spline snake is evolved to refine the leukocytes boundaries by maximizing the associated GICOV score. The third and final step retains only the extracted contours that have a GICOV score above the analytically determined threshold. Experimental results using 327 rolling leukocytes were compared to those of human experts and currently used methods. The proposed GICOV method achieves 78.6% leukocyte detection accuracy with 13.1% false alarm rate.
Gang Dong, Nilanjan Ray, Scott T. Acton
IEEE Trans. Medical Imaging1
2004 Detection of Microspheres in Venules for Automated Particle Image
abstract
In this paper, we propose an automatic approach for detecting particle tracers (microspheres) in microscopic imagery obtained from mouse cremaster venules in vivo. Measurements of the translational speed and radial position of individual microspheres provide the input data needed to extract velocity profiles from steady blood flow in venules. These profiles provide information about local hemodynamics that is critical to a broad range of fields in microvascular physiology, including endothelial-cell mechanotransduction, inflammation, and microvascular resistance. In the preprocessing stage, an active contour method based on dynamic programming is used for vessel region extraction. Each microsphere is then identified using a process of coarse segmentation followed by verification. Segmentation is achieved using a morphological method for microsphere detection while verification is achieved using an analytical model tailored to the microsphere. Experimental results are obtained using the proposed scheme and compared with previously published manually acquired data.
Gang Dong, Edward Damiano, Michael L. Smith, Scott T. Acton, Klaus Ley
CBMS1
2003 A variational method for leukocyte detection
abstract
In this paper, we propose a variational method for the detection of leukocytes observed in vivo. An adaptive threshold surface is constructed automatically using boundary information from the image. The surface is created using an objective functional that is minimized via a variational approach. This surface is constrained by an edge field that is also computed with a variational method. Objects extracted from background are pruned according to two geometric criteria. In the experiments, we find the false positive rate of the detector and show that the proposed approach can automatically and accurately identify multiple rolling leukocytes in vivo.
Gang Dong, Scott T. Acton
ICIP (2)1