Yiming Du

dblp:230/6249 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents
abstract
Modern task-oriented dialogue (TOD) systems increasingly rely on large language model (LLM) agents, leveraging Retrieval-Augmented Generation (RAG) and long-context capabilities for long-term memory utilization. However, these methods prioritise semantic similarity over task intent, degrading multi-session coherence. We propose MemGuide, a two-stage intent-driven memory selection framework: (1) Intent‑Aligned Retrieval retrieves goal-consistent QA‑formatted memory units; (2) Missing‑Slot Guided Filtering reranks units by slot-completion gain via a chain‑of‑thought reasoner and fine‑tuned LLaMA‑8B filter. We also introduce the MS-TOD, the first multi-session TOD benchmark with 132 diverse personas, 956 task goals, and annotated intent-aligned memory targets. Evaluations on MS-TOD show that MemGuide boosts task success rate by 11% (88%→99%) and reduces dialogue length by 2.84 turns, and matches single‑session performance.
Yiming Du, Bin Liang 0004, Baojun Wang, Lin Gui 0003, Jeff Z. Pan, Ruifeng Xu 0001, Kam-Fai Wong
AAAI1
2026 Deep Incomplete Multi-View Clustering via Hierarchical Imputation and Alignment
abstract
Incomplete multi-view clustering (IMVC) aims to discover shared cluster structures from multi-view data with partial observations. The core challenges lie in accurately imputing missing views without introducing bias, while maintaining semantic consistency across views and compactness within clusters. To address these challenges, we propose DIMVC-HIA, a novel deep IMVC framework that integrates hierarchical imputation and alignment with four key components: (1) view-specific autoencoders for latent feature extraction, coupled with a view-shared clustering predictor to produce soft cluster assignments; (2) a hierarchical imputation module that first estimates missing cluster assignments based on cross-view contrastive similarity, and then reconstructs missing features using intra-view, intra-cluster statistics; (3) an energy-based semantic alignment module, which promotes intra-cluster compactness by minimizing energy variance around low-energy cluster anchors; and (4) a contrastive assignment alignment module, which enhances cross-view consistency and encourages confident, well-separated cluster predictions. Experiments on benchmarks demonstrate that our framework achieves superior performance under varying levels of missingness.
Yiming Du, Rui Ning, Lusi Li
AAAI1
2026 Mitigating Context Interference for Reliable and Efficient Search Agents
abstract
Boyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boyang Xue, Bin Wu 0025, Shuofei Qiao, Rui Wang 0092, Yiming Du, Hongru Wang 0003, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani
ACL (1)6
2026 EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue Systems
abstract
Zhengyi Zhao, Shubo Zhang, Yiming Du, Bin Liang, Baojun Wang, Zhongyang Li, Binyang Li, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengyi Zhao 0001, Shubo Zhang, Yiming Du, Bin Liang 0004, Baojun Wang, Binyang Li, Kam-Fai Wong
ACL (1)3
2026 CHIP-MAP: A Collaborative Optimization Framework for Macro Placement Using Large Language Models
abstract
As integrated circuits continue to grow in both scale and complexity, macro placement plays a critical role in physical design, directly affecting chip-level performance, power, and area (PPA). Traditional macro placement methods, such as simulated annealing, analytical optimization, and reinforcement learning, face limitations including slow convergence, heavy dependence on large datasets, and over-reliance on intermediate PPA indicators rather than final PPA. Large language models (LLMs) offer strong generative power and semantic reasoning that can potentially automate macro layout tasks while addressing the aforementioned problems in traditional methods, but their limited understanding of layout rules and lack of iterative, feedback-driven refinement make direct application challenging. To address this, we propose CHIP-MAP, a macro placement framework based on multi-agent collaboration and feedback-driven optimization. Furthermore, we introduce two innovative tools: the Module Link Weight Analyzer (MWA) and the Standard Cell Usability Score (SCUS), which are designed to guide fine-grained layout refinement. We evaluate CHIP-MAP on five benchmarks ranging from low-power cores to large multi-core processors implemented at 130nm and 45nm technology nodes. Results show that it achieves up to 1.5% area reduction and an average repair of 61.6% of total negative slack (TNS), while also reducing wirelength and improving timing.
Yiming Du, Renye Yan, Yunfan Yang, Frank Qu, Jiajun Tan, ZhiYu Zheng, Yiming Gan, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai
DATE1
2026 SONIC: Smart Optimization for Neural-Integrated CMP with Timing-Aware Fills
abstract
Dummy fill insertion is essential for CMP uniformity but remains challenging due to the nonlinear CMP process, the large optimization space, and timing degradation caused by parasitic coupling. We propose SONIC, a differentiable CMP-driven dummy fill optimization framework that employs a neural CMP simulator to directly optimize planarization objectives using gradient-based methods. SONIC further integrates a timing-aware fill insertion strategy to mitigate coupling capacitance near critical nets. Experimental results demonstrate that SONIC achieves competitive planarization quality with up to 1830× runtime speedup over a full-chip CMP simulator. Compared with the state-of-the-art model-based method, SONIC reduces height variation, line deviation, and outliers by up to 86.16%, 90.10%, and 51.61%, respectively, while achieving a 77.67% runtime reduction and lowering coupling capacitance by 13.05%.
Jiajun Tan, Yiming Du, Yiming Gan, Ling Lang 0002, Yibo Lin, Zongwei Wang 0001, Yimao Cai
DATE3
2026 PGFormer: A Prototype-Graph Transformer for Incomplete Multiview Clustering
abstract
Incomplete multiview clustering (IMVC) faces significant challenges due to missing data and inherent view discrepancies. While deep neural networks offer powerful representation learning capabilities for IMVC, existing methods often overlook view diversity and force representations across views to be identical, leading to 1) biased representations with distorted topologies and 2) inaccurate imputation for missing data, ultimately degrading clustering performance. To address these issues, we propose prototype-graph transformer (PGFormer), a novel IMVC framework that integrates prototype assignments, rather than direct representations, to enhance clustering performance. PGFormer leverages view-specific encoders to extract features from available samples in each view, employs a PGFormer designed to refine node embeddings, and reconstructs available samples using these refined embeddings. For each view, PGFormer utilizes a graph convolutional network (GCN) to model node-to-node topologies and generate semantic prototypes from the node embeddings. These view-specific prototypes and embeddings are then refined through dual attention mechanisms: prototype-to-prototype (P2P) self-attention and prototype-to-node (P2N) cross-attention, enabling a thorough exploration of multilevel topological relationships within each view. To address missing data, the cross-prototype imputation (CPI) module leverages the weighted prototype assignments from different views to impute missing samples using refined intraview prototypes. Building on this, the cross-view alignment module calibrates prototype assignments to ensure consistent predictions across views. Extensive experiments demonstrate that PGFormer can achieve superior performance compared with the baselines.
Yiming Du, Rui Ning, Lusi Li
IEEE Trans. Neural Networks Learn. Syst.1
2025 A New Formula for Sticker Retrieval: Reply with Stickers in Multi-Modal and Multi-Session Conversation
abstract
Stickers are widely used in online chatting, which can vividly express someone's intention, emotion, or attitude. Existing conversation research typically retrieves stickers based on a single session or the previous textual information, which can not adapt to the multi-modal and multi-session nature of the real-world conversation. To this end, we introduce MultiChat, a new dataset for sticker retrieval facing the multi-modal and multi-session conversation, comprising 1,542 sessions, featuring 50,192 utterances and 2,182 stickers. Based on the created dataset, we propose a novel Intent-Guided Sticker Retrieval (IGSR) framework that retrieves stickers for multi-modal and multi-session conversation history drawing support from intent learning. Specifically, we introduce sticker attributes to better leverage the sticker information in multi-modal conversation, which are incorporated with utterances to construct a memory bank. Further, we extract relevant memories for the current conversation from the memory bank to identify the intent of the current conversation, and then retrieve a sticker to respond guided by the intent. Extensive experiments on our MultiChat dataset reveal the robustness and effectiveness of our IGSR approach in multi-session, multi-modal scenarios.
Yiming Du, Bin Liang 0004, Zhixin Bai, Min Yang 0007, Baojun Wang, Kam-Fai Wong, Ruifeng Xu 0001
AAAI2
2025 ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
abstract
Fine-tuning multi-turn dialogue systems requires high-quality supervision but often suffers from degraded performance when exposed to low-quality data.Supervision errors in early turns can propagate across subsequent turns, undermining coherence and response quality.Existing methods typically address data quality via static prefiltering, which decouples quality control from training and fails to mitigate turn-level error propagation.In this context, we propose ReSURE (Regularizing Supervision UnREliability), an adaptive learning method that dynamically downweights unreliable supervision without explicit filtering.ReSURE estimates per-turn loss distributions using Welford's online statistics and reweights sample losses on the fly accordingly.Experiments on both singlesource and mixed-quality datasets show improved stability and response quality.Notably, ReSURE enjoys positive Spearman correlations (0.21 ∼ 1.0 across multiple benchmarks) between response scores and number of samples regardless of data quality, which potentially paves the way for utilizing large-scale data effectively.
Yiming Du, Yifan Xiang, Bin Liang 0004, Dahua Lin, Kam-Fai Wong, Fei Tan 0002
EMNLP1
2025 Flexibly Utilize Memory for Long-Term Conversation via a Fragment-then-Compose Framework
abstract
Cai Ke, Yiming Du, Bin Liang, Yifan Xiang, Lin Gui, Zhongyang Li, Baojun Wang, Yue Yu, Hui Wang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Cai Ke, Yiming Du, Bin Liang 0004, Yifan Xiang, Lin Gui 0003, Baojun Wang, Yue Yu 0001, Hui Wang 0030, Kam-Fai Wong, Ruifeng Xu 0001
EMNLP2
2025 Energy-based Deep Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) deals with real-world scenarios where certain views are partially missing, posing significant challenges to effective clustering. Most existing IMVC approaches face a trade-off: imputation-free methods suffer from information bias and imbalance, while full-imputation methods risk introducing and propagating noise. To overcome these limitations, we propose Energy-Based Deep Incomplete Multi-View Clustering (Energy-DIMC), a novel selective-imputation framework that leverages energy-based models (EBMs) to guide reliable imputations and robust clustering. EBMs assess data compatibility by assigning lower energy to more coherent structures, effectively modeling complex inter-view and inter-sample dependencies. Inspired by EBMs, Energy-DIMC integrates four key components: 1) a view feature projector that learns view-specific features and projects them into a common feature space; 2) an energy-guided selective imputation module that identifies the most reliable source view for each view based on view energies, and performs feature imputation only when cross-view transfer is feasible, avoiding unreliable imputations; 3) an energy-based representation fusion module that aggregates observed and selectively imputed features across views via a view attention mechanism, generating view-coherent representations; 4) an energy-enhanced contrastive alignment module that enforces consistency between view-specific and view-coherent representations using dual-level energy signals to preserve true positives. Extensive experiments demonstrate that Energy-DIMC outperforms state-of-the-art IMVC methods across diverse missing-view scenarios. The code is available at https://github.com/sunway677/EnergyIMVC.
Yiming Du, Rui Ning, Lusi Li
ACM Multimedia2
2025 Deep Content and Contrastive Perception learning for automatic fetal nuchal translucency image quality assessment
Weiping Ding 0001, Jinzhao Yang, Huiyu Zhou 0001, Yiming Du, Bin Hu 0023, Lichi Zhang, Qian Wang 0001
Eng. Appl. Artif. Intell.7
2025 Deep Incomplete Multi-view Clustering via Multi-level Imputation and Contrastive Alignment
Yiming Du, Rui Ning, Lusi Li
Neural Networks2
2024 M3sum: A Novel Unsupervised Language-Guided Video Summarization
abstract
Language-guided video summarization empowers users to use natural language queries to effortlessly summarize lengthy videos into concise and relevant summaries that cater specifically to their information needs, which is more friendly to access and digest. However, most of the previous works rely on tremendous (also expensive) annotated videos and complex designs to align different modals at the feature level. In this paper, we first explore the combination of off-the-shelf models for each modal to solve the complex multi-modal problem by proposing a novel unsupervised language-guided video summarization method: Modular Multi-Modal Summarization (M3Sum), which does not require any training data or parameter updates. Specifically, instead of training an alignment module at the feature level, we convert all modal information (e.g. audio and frames) into textual descriptions and design a parameter-free alignment mechanism to fuse text descriptions from different modals. Benefiting from the remarkable long-context understanding capability of large language models (LLMs), our approach demonstrates comparable performance to most unsupervised methods and even outperforms certain supervised methods.
Hongru Wang 0003, Baohang Zhou, Zhengkun Zhang, Yiming Du, David Ho, Kam-Fai Wong
ICASSP4
2023 HugeGPT: Storing Guest Page Tables on Host Huge Pages to Accelerate Address Translation
abstract
Expensive page table walks triggered by frequent TLB misses have incurred major performance bottlenecks for data-intensive workloads that are dominated by memory accesses with weak locality. Since it is hard to reduce TLB misses for such workloads, reducing page table walk overhead (i.e., the overhead of each TLB miss) is an increasingly important direction for improving application performance. The direction is more compelling for workloads running in virtual machines (VMs). In virtualized environments, each TLB miss triggers a two-dimensional page table walk, which has a significantly higher overhead than that on native systems. This paper presents HugeGPT, a software approach to reducing two-dimensional page table walk overhead in virtualized environments. HugeGPT ensures that page tables used in guest systems are physically held in the huge pages formed in the host. This brings two-fold benefits: 1) the number of steps walking down the host page table is reduced; 2) the misses of page walk caches incurred by accessing the leaf nodes on host page tables can be eliminated. Extensive evaluation based on the prototype implementation and diverse real-world applications shows that HugeGPT can efficiently reduce address translation overhead and improve application performance in virtualized clouds.
Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Yiming Du, Xiaoning Ding, Tianyin Xu
PACT4
2022 Grey wolf optimizer based on Aquila exploration method
Haisong Huang, Qingsong Fan, Jianan Wei, Yiming Du, Weisen Gao
Expert Syst. Appl.5
2020 Channel Estimation and Transmission for Intelligent Reflecting Surface Assisted THz Communications
abstract
Intelligent reflecting surface (IRS) is envisioned as a promising technology to broaden signal coverage and enhance transmission in terahertz (THz) communications. Due to the passivity of IRS, the channel measurement can not be achieved by traditional pilot manner and the subsequent cooperative transmission design remains an open problem. This paper investigates the channel estimation and transmission solutions for massive multiple input multiple output (MIMO) IRS-assisted THz system. The channel estimation is realized by beam training and the quantization error is analyzed for evaluating performance. In addition, a novel hierarchical search codebook design is proposed as a low-complexity basis of beam training. Based on above foundations, we propose a cooperative channel estimation procedure to tactfully acquire the channel knowledge. Finally, by leveraging obtained channel information, the designs of IRS and transceivers are directly provided in closed form without reconstructing the full channel matrix or additional optimization. Simulation and numerical results are presented to illustrate the minimum signal to noise ratio (SNR) required for beam training and the efficacy of the proposed transmission solutions.
Boyu Ning, Zhi Chen 0002, Wenrong Chen, Yiming Du
ICC4
2019 Motor imagery EEG recognition with KNN-based smooth auto-encoder
Yiming Du, Yuyan Dai
Artif. Intell. Medicine3