Shuang Zeng

dblp:266/1045 · DBLP profile ↗
← Back
40ranked-venue papers
13as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 8 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
abstract
Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture their persistent effectiveness across extended driving sequences. In this paper, we present PAMR (Persistent Autoregressive Mapping with Traffic Rules), a novel framework that performs autoregressive co-construction of lane vectors and traffic rules from visual observations. Our approach introduces two key mechanisms: Map-Rule Co-Construction for processing driving scenes in temporal segments, and Map-Rule Cache for maintaining rule consistency across these segments. To properly evaluate continuous and consistent map generation, we develop MapDRv2, featuring improved lane geometry annotations. Extensive experiments demonstrate that PAMR achieves superior performance in joint vector-rule mapping tasks, while maintaining persistent rule effectiveness throughout extended driving sequences.
Shiyi Liang, Xinyuan Chang, Changjie Wu, Huiyuan Yan, Yifan Bai 0001, Yujian Yuan, Shuang Zeng, Mu Xu, Xing Wei 0001
AAAI9
2026 Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
abstract
Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on complex reasoning tasks. However, the one-sided reward, focused solely on final correctness, limits its ability to provide detailed supervision over internal reasoning process. This deficiency leads to suboptimal internal reasoning quality, manifesting as issues like over-thinking, under-thinking, redundant-thinking, and disordered-thinking. Inspired by the recent progress in LRM self-rewarding, we introduce self-rewriting framework, where a model rewrites its own reasoning texts, and subsequently learns from the rewritten reasoning to improve the internal thought process quality. For algorithm design, we propose a selective rewriting approach wherein only "simple" samples, defined by the model's consistent correctness, are rewritten, thereby preserving all original reward signals of GRPO. For practical implementation, we compile rewriting and vanilla generation within one single batch, maintaining the scalability of the RL algorithm and introducing only 10% overhead. Extensive experiments on diverse tasks with different model sizes validate the effectiveness of self-rewriting. In terms of the accuracy-length tradeoff, the self-rewriting approach achieves improved accuracy (+0.6) with substantially shorter reasoning (-46%) even without explicit instructions in rewriting prompts to reduce reasoning length, outperforming existing strong baselines. In terms of internal reasoning quality, self-rewriting achieves significantly higher scores (+7.2) under the LLM-as-a-judge metric, successfully mitigating internal reasoning flaws.
Jiashu Yao, Heyan Huang, Shuang Zeng, Chuwei Luo, WangJie You, Jie Tang 0001, Yuhang Guo 0001, Yangyang Kang
AAAI3
2026 UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
abstract
Large-scale map construction is foundational for critical applications such as autonomous driving and navigation systems. Traditional large-scale map construction approaches mainly rely on costly and inefficient special data collection vehicles and labor-intensive annotation processes. While existing satellite-based methods have demonstrated promising potential in enhancing the efficiency and coverage of map construction, they exhibit two major limitations: (1) inherent drawbacks of satellite data (e.g., occlusions, outdatedness) and (2) inefficient vectorization from perception-based methods, resulting in discontinuous and rough roads that require extensive post-processing. This paper presents a novel generative framework, UniMapGen, for large-scale map construction, offering three key innovations: (1) representing lane lines as discrete sequence and establishing an iterative strategy to generate more complete and smooth map vectors than traditional perception-based methods. (2) proposing a flexible architecture that supports multi-modal inputs, enabling dynamic selection among BEV, PV, and text prompt, to overcome the drawbacks of satellite data. (3) developing a state update strategy for global continuity and consistency of the constructed large-scale map. UniMapGen achieves state-of-the-art performance on the OpenSatMap dataset. Furthermore, UniMapGen can infer occluded roads and predict roads missing from dataset annotations.
Yujian Yuan, Changjie Wu, Xinyuan Chang, Sijin Wang, Shiyi Liang, Shuang Zeng, Mu Xu
AAAI7
2026 PriorDrive: Enhancing Online HD Mapping with Unified Vector Priors
abstract
High-Definition Maps (HD maps) are essential for the precise navigation and decision-making of autonomous vehicles, yet their creation and upkeep present significant cost and timeliness challenges. The online construction of HD maps using on-board sensors has emerged as a promising solution; however, these methods can be impeded by incomplete data due to occlusions and inclement weather, while their performance in distant regions remains unsatisfying. This paper proposes PriorDrive to address these limitations by directly harnessing the power of various vectorized prior maps, significantly enhancing the robustness and accuracy of online HD map construction. Our approach integrates a variety of prior maps uniformly, such as OpenStreetMap's Standard Definition Maps (SD maps), outdated HD maps from vendors, and locally constructed maps from historical vehicle data. To effectively integrate such prior information into online mapping models, we introduce a Hybrid Prior Representation (HPQuery) that standardizes the representation of diverse map elements. We further propose a Unified Vector Encoder (UVE), which employs fused prior embedding and a dual encoding mechanism to encode vector data. To improve the UVE's generalizability and performance, we propose a segment-level and point-level pre-training strategy that enables the UVE to learn the prior distribution of vector data. Through extensive testing on the nuScenes, Argoverse 2 and OpenLane-V2, we demonstrate that PriorDrive is highly compatible with various online mapping models and substantially improves map prediction capabilities. The integration of prior maps through PriorDrive offers a robust solution to the challenges of single-perception data, paving the way for more reliable autonomous driving.
Shuang Zeng, Xinyuan Chang, Yujian Yuan, Shiyi Liang, Mu Xu, Xing Wei 0001
AAAI1
2026 MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics
abstract
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May Dongmei Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng 0006, Shuang Zeng, Jinzhuo Wang, May D. Wang
ACL (1)6
2026 MPRL: Multi-Perspective Reinforcement Learning for Enhancing Format Adherence Capability of Large Language Models
Shuang Zeng, Qiaochen Wang
PAKDD (1)3
2026 Improve retinal artery/vein classification via channel coupling
Shuang Zeng, Chee Hong Lee, Boxu Xie, Ourui Fu, Hangzhou He, Lei Zhu 0012, Yanye Lu, Fangxiao Cheng
Expert Syst. Appl.1
2026 Stochastic Gradient Methods: Bias, Stability and Generalization
abstract
Recent developments of stochastic optimization often suggest biased gradient estimators to improve either the robustness, communication efficiency or computational speed. Representative biased stochastic gradient methods (BSGMs) include Zeroth-order stochastic gradient descent (SGD), Clipped-SGD and SGD with delayed gradients. The practical success of BSGMs motivates a lot of convergence analysis to explain their impressive training behaviour. As a comparison, there is far less work on their generalization analysis, which is a central topic in modern machine learning. In this paper, we present the first framework to study the stability and generalization of BSGMs for convex and smooth problems. We introduce a generalized Lipschitz-type condition on gradient estimators and bias, under which we develop a rather general stability bound to show how the bias and the gradient estimators affect the stability. We apply our general result to develop the first stability bound for Zeroth-order SGD with reasonable step size sequences, and the first stability bound for Clipped-SGD. While our stability analysis is developed for general BSGMs, the resulting stability bounds for both Zeroth-order SGD and Clipped-SGD match those of SGD under appropriate smoothing/clipping parameters. We combine the stability and convergence analysis together, and derive excess risk bounds of order $O(1/\sqrt{n})$ for both Zeroth-order SGD and Clipped-SGD, where $n$ is the sample size.
Shuang Zeng, Yunwen Lei
J. Mach. Learn. Res.1
2026 CASLO: Joint scaling and deployment for microservices leveraging context-aware SLO assignment
Shuang Zeng, Zezhong Yan
J. Netw. Comput. Appl.1
2026 Exploring the Vulnerabilities of Federated Learning: A Deep Dive Into Gradient Inversion Attacks
abstract
Federated Learning (FL) has emerged as a promising privacy-preserving collaborative model training paradigm without sharing raw data. However, recent studies have revealed that private information can still be leaked through shared gradient information and attacked by Gradient Inversion Attacks (GIA). While many GIA methods have been proposed, a detailed analysis, evaluation, and summary of these methods are still lacking. Although various survey papers summarize existing privacy attacks in FL, few studies have conducted extensive experiments to unveil the effectiveness of GIA and their associated limiting factors in this context. To fill this gap, we first undertake a systematic review of GIA and categorize existing methods into three types, i.e., optimization-based GIA (OP-GIA), generation-based GIA (GEN-GIA), and analytics-based GIA (ANA-GIA). Then, we comprehensively analyze and evaluate the three types of GIA in FL, providing insights into the factors that influence their performance, practicality, and potential threats. Our findings indicate that OP-GIA is the most practical attack setting despite its unsatisfactory performance, while GEN-GIA has many dependencies and ANA-GIA is easily detectable, making them both impractical. Finally, we offer a three-stage defense pipeline to users when designing FL frameworks and protocols for better privacy protection and share some future research directions from the perspectives of attackers and defenders that we believe should be pursued. We hope that our study can help researchers design more robust FL frameworks to defend against these attacks.
Pengxin Guo 0001, Runxi Wang, Shuang Zeng, Jinjing Zhu, Haoning Jiang, Yuyin Zhou, Hui Xiong 0001, Liangqiong Qu
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 HPCTrans: Heterogeneous Plumage Cues-Aware Texton Correlation Representation for FBIC via Transformers
abstract
Fine-grained bird image classification (FBIC) for distinguishing bird subspecies is challenging because of several issues, including a camouflaged appearance, body occlusion, and an arbitrary bird posture. To address these challenges, we propose a novel heterogeneous plumage cues-aware texton correlation representation for FBIC, which leverages texton correlation in various functional plumage regions for effective learning. Two key findings are revealed: 1) texton structural discrepancies of heterogeneous plumage; and 2) abstract region information for specific birds. On this basis, this model introduces texton coherence extraction module (TCEM) and abstract representation selection (ARS). Specifically, considering bird characteristics, TCEM is introduced to exploit the spatial statistical properties of local textons in heterogeneous plumage. To the best of our knowledge, this study is the first to introduce heterogeneous plumage cues for mining texton correlation relationship representations in FBIC tasks. In addition, a Multiscale Information Cross-Attention Transformer (MICAformer) is proposed for better modeling texton correlation representation. The experimental results on the CUB-200-2011 dataset and NABirds show the effectiveness of the proposed HPCTrans model over the state-of-the-art methods.
Hai Liu 0004, Shuang Zeng, Liqian Deng, Tingting Liu 0006, Xionghua Liu, Zhaoli Zhang, Youfu Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-Training
abstract
Medical image segmentation is a critical yet challenging task, primarily due to the difficulty of obtaining extensive datasets of high-quality, expert-annotated images. Contrastive learning presents a potential but still problematic solution to this issue. Because most existing methods focus on extracting instance-level or pixel-to-pixel representation, which ignores the characteristics between intra-image similar pixel groups. Moreover, when considering contrastive pairs generation, most SOTA methods mainly rely on manually setting thresholds, which requires a large number of gradient experiments and lacks efficiency and generalization. To address these issues, we propose a novel contrastive learning approach named SuperCL for medical image segmentation pre-training. Specifically, our SuperCL exploits the structural prior and pixel correlation of images by introducing two novel contrastive pairs generation strategies: Intra-image Local Contrastive Pairs (ILCP) Generation and Inter-image Global Contrastive Pairs (IGCP) Generation. Considering superpixel cluster aligns well with the concept of contrastive pairs generation, we utilize the superpixel map to generate pseudo masks for both ILCP and IGCP to guide supervised contrastive learning. Moreover, we also propose two modules named Average SuperPixel Feature Map Generation (ASP) and Connected Components Label Generation (CCL) to better exploit the prior structural information for IGCP. Finally, experiments on 8 medical image datasets indicate our SuperCL outperforms existing 12 methods. i.e. Our SuperCL achieves a superior performance with more precise predictions from visualization figures and 3.15%, 5.44%, 7.89% DSC higher than the previous best results on MMWHS, CHAOS, Spleen with 10% annotations. Our code is released at https://github.com/stevezs315/SuperCL.
Shuang Zeng, Lei Zhu 0012, Hangzhou He, Yanye Lu
IEEE Trans. Image Process.1
2025 A New Federated Learning Framework Against Gradient Inversion Attacks
abstract
Federated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety of privacy-preserving methods have been integrated into FL to thwart such attacks, such as Secure Multi-party Computing (SMC), Homomorphic Encryption (HE), and Differential Privacy (DP). Despite their ability to protect data privacy, these approaches inherently involve substantial privacy-utility trade-offs. By revisiting the key to privacy exposure in FL under GIA, which lies in the frequent sharing of model gradients that contain private data, we take a new perspective by designing a novel privacy preserve FL framework that effectively ``breaks the direct connection'' between the shared parameters and the local private data to defend against GIA. Specifically, we propose a Hypernetwork Federated Learning (HyperFL) framework that utilizes hypernetworks to generate the parameters of the local model and only the hypernetwork parameters are uploaded to the server for aggregation. Theoretical analyses demonstrate the convergence rate of the proposed HyperFL, while extensive experimental results show the privacy-preserving capability and comparable performance of HyperFL.
Pengxin Guo 0001, Shuang Zeng, Xiaodan Zhang 0003, Weihong Ren, Yuyin Zhou, Liangqiong Qu
AAAI2
2025 V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer
abstract
Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. Recent approaches have leveraged the knowledge of large language models to construct concept bottlenecks, with multimodal models like CLIP subsequently mapping image features into the concept feature space for classification. Despite this, the concepts produced by language models can be verbose and may introduce non-visual attributes, which hurts accuracy and interpretability. In this study, we investigate to avoid these issues by constructing CBMs directly from multimodal models. To this end, we adopt common words as base concept vocabulary and leverage auxiliary unlabeled images to construct a Vision-to-Concept (V2C) tokenizer that can explicitly quantize images into their most relevant visual concepts, thus creating a vision-oriented concept bottleneck tightly coupled with the multimodal model. This leads to our V2C-CBM which is training efficient and interpretable with high accuracy. Our V2C-CBM has matched or outperformed LLM-supervised CBMs on various visual classification benchmarks, validating the efficacy of our approach.
Hangzhou He, Lei Zhu 0012, Shuang Zeng, Yanye Lu
AAAI4
2025 CoMPI: Coordinated Model Merging and Parallel Inference at Edge
abstract
Modern edge intelligence applications increasingly rely on multiple deep learning models operating under strict latency and memory constraints. However, the limited GPU memory available in edge clusters often leads to costly model switching and frequent SLO violations, particularly under bursty workloads. Several techniques are effective in reducing memory pressure from different aspects: model merging reduces memory usage on individual GPUs by sharing parameters across models, while parallelism reduces memory usage across GPUs by reducing redundant replicas. Since these two techniques operate in orthogonal decision spaces, combining them in edge clusters offers the potential for greater memory efficiency. However, real-world deployments reveal implicit interdependen-cies between them, making independent optimization inadequate for realizing their combined advantages.
Shuang Zeng, Zezhong Yan
SoCC1
2025 SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions
abstract
Accurate lane topology is essential for autonomous driving, yet traditional methods struggle to model the complex, non-linear structures-such as loops and bidirectional lanes-prevalent in real-world road structure. We present SeqGrowGraph, a novel framework that learns lane topology as a chain of graph expansions, inspired by human map-drawing processes. Representing the lane graph as a directed graph $G=(V,E)$, with intersections ($V$) and centerlines ($E$), SeqGrowGraph incrementally constructs this graph by introducing one vertex at a time. At each step, an adjacency matrix ($A$) expands from $n \times n$ to $(n+1) \times (n+1)$ to encode connectivity, while a geometric matrix ($M$) captures centerline shapes as quadratic Bézier curves. The graph is serialized into sequences, enabling a transformer model to autoregressively predict the chain of expansions, guided by a depth-first search ordering. Evaluated on nuScenes and Argoverse 2 datasets, SeqGrowGraph achieves state-of-the-art performance.
Mengwei Xie, Shuang Zeng, Xinyuan Chang, Mu Xu, Xing Wei 0001
ICCV2
2025 Selective Aggregation for Low-Rank Adaptation in Federated Learning
abstract
We investigate LoRA in federated learning through the lens of the asymmetry analysis of the learned $A$ and $B$ matrices. In doing so, we uncover that $A$ matrices are responsible for learning general knowledge, while $B$ matrices focus on capturing client-specific knowledge. Based on this finding, we introduce Federated Share-A Low-Rank Adaptation (FedSA-LoRA), which employs two low-rank trainable matrices $A$ and $B$ to model the weight update, but only $A$ matrices are shared with the server for aggregation. Moreover, we delve into the relationship between the learned $A$ and $B$ matrices in other LoRA variants, such as rsLoRA and VeRA, revealing a consistent pattern. Consequently, we extend our FedSA-LoRA method to these LoRA variants, resulting in FedSA-rsLoRA and FedSA-VeRA. In this way, we establish a general paradigm for integrating LoRA with FL, offering guidance for future work on subsequent LoRA variants combined with FL. Extensive experimental results on natural language understanding and generation tasks demonstrate the effectiveness of the proposed method. Our code is available at https://github.com/Pengxin-Guo/FedSA-LoRA.
Pengxin Guo 0001, Shuang Zeng, Huijie Fan, Liangqiong Qu
ICLR2
2025 Stability and Generalization Analysis of Decentralized SGD: Sharper Bounds Beyond Lipschitzness and Smoothness
abstract
Decentralized SGD (D-SGD) is a popular optimization method to train large-scale machine learning models. In this paper, we study the generalization behavior of D-SGD for both smooth and nonsmooth problems by leveraging the algorithm stability. For convex and smooth problems, we develop stability bounds involving the training errors to show the benefit of optimization in generalization. This improves the existing results by removing the Lipschitzness assumption and implying fast rates in a low-noise condition. We also develop the first optimal stability-based generalization bounds for D-SGD applied to nonsmooth problems. We further develop optimization error bounds which imply minimax optimal excess risk rates. Our novelty in the analysis consists of an error decomposition to use the co-coercivity of functions as well as the control of a neighboring-consensus error.
Shuang Zeng, Yunwen Lei
ICML1
2025 FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
abstract
Vision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability. Most prior work, however, inserts textual chains-of-thought (CoT) as intermediate steps tailored to the current scene. Such symbolic compressions can blur spatio-temporal relations and discard fine visual cues, creating a cross-modal gap between perception and planning. We propose FSDrive, a visual spatio-temporal CoT framework that enables VLAs to think in images. The model first acts as a world model to generate a unified future frame that overlays coarse but physically-plausible priors—future lane dividers and 3D boxes—on the predicted future image. This unified frame serves as the visual CoT, capturing both spatial structure and temporal evolution. The same VLA then functions as an inverse-dynamics model, planning trajectories from current observations and the visual CoT. To equip VLAs with image generation while preserving understanding, we introduce a unified pre-training paradigm that expands the vocabulary to include visual tokens and jointly optimizes VQA (for semantics) and future-frame prediction (for dynamics). A progressive easy-to-hard scheme first predicts lane/box priors to enforce physical constraints, then completes full future frames for fine details. On nuScenes and NAVSIM, FSDrive improves trajectory accuracy and reduces collisions under both ST-P3 and UniAD metrics, and attains competitive FID for future-frame generation despite using lightweight autoregression. It also advances scene understanding on DriveLM. Together, these results indicate that visual CoT narrows the cross-modal gap and yields safer, more anticipatory planning. Code is available at https://github.com/MIV-XJTU/FSDrive.
Shuang Zeng, Xinyuan Chang, Mengwei Xie, Yifan Bai 0001, Mu Xu, Xing Wei 0001
NeurIPS1
2025 Exploiting module evolution correlation relationship for fine-grained bird image classification with structural functional representation
Shuang Zeng, Hai Liu 0004, Tingting Liu 0006, Qiuxia Liu, Minhong Wang 0001, Zhaoli Zhang
Neurocomputing1
2025 Novel extraction of discriminative fine-grained feature to improve retinal vessel segmentation
Shuang Zeng, Chee Hong Lee, Micky C. Nnamdi, Wenqi Shi 0002, J. Ben Tamo, Hangzhou He, May D. Wang, Lei Zhu 0012, Yanye Lu, Qiushi Ren
Image Vis. Comput.1
2025 Points-Supervised Fundus Vessel Segmentation via Shape Priors and Contrastive Learning
abstract
The performance of fully supervised methods for fundus vessel segmentation highly relies on a large number of full labels which are laborious and time-consuming to obtain. Although weak annotations relax the requirement for pixel-wise labeling, they pose challenges in learning comprehensive information about the target. Some methods use pseudo labels generated from network predictions for extra supervision, but false positive predictions in these labels may harm training. In this paper, to tackle this problem and to balance the annotation cost and supervision information, we introduce point annotations to fundus vessel segmentation and propose a novel method, called Points-based Vessel segmentation Network (PVN), to enhance the segmentation accuracy. In PVN, to avoid noise in pseudo labels, by combining proposed Point Activation Maps, shape priors of vessels are learned and used as soft supervision. Additionally, to further leverage the annotated vessel and background points, we design a novel contrastive learning method in a pixels-and-regions-mixed manner, which helps learn discriminative features by distinguishing between pixel and region samples of vessels and background. The performance of PVN is evaluated on laser speckle contrast imaging fundus images, 548 nm fundus images, and three public datasets, where PVN outperforms other point-supervised methods. Even with only 1% annotated pixels, PVN still achieves excellent performance. Our method is also flexible and easy to be combined with other frameworks. To the best of our knowledge, we are the first to propose and demonstrate the effectiveness of point annotations for fundus vessel segmentation. Our code is available at: https://github.com/kaiwenli325/PVN.
Hangzhou He, Shuang Zeng, Lei Zhu 0012, Yanye Lu
IEEE Trans. Medical Imaging3
2025 Branches Mutual Promotion for End-to-End Weakly Supervised Semantic Segmentation
abstract
End-to-end weakly supervised semantic segmentation (E2E-WSSS) aims at optimizing a segmentation model in a single-stage training process based on only image annotations. Existing methods adopt an online-trained classification branch to provide pseudo annotations for supervising the segmentation branch. However, this strategy makes the classification branch dominate the whole concurrent training process, hindering these two branches from assisting each other. In our work, we treat these two branches equally by viewing them as diverse ways to generate the segmentation map, and add interactions on both their supervision and operation to achieve mutual promotion. For this purpose, a bidirectional supervision mechanism is elaborated to force the consistency between the outputs of these two branches. Thus, the segmentation branch can also give feedback to the classification branch to enhance the quality of localization seeds. Moreover, our method also designs interaction operations between these two branches to exchange their knowledge to assist each other. Experiments indicate our work outperforms existing end-to-end weakly supervised segmentation methods. Codes are available at https://github.com/zh460045050/BMP-WSSS.
Lei Zhu 0012, Hangzhou He, Shuang Zeng, Yibao Zhang, Qiushi Ren, Yanye Lu
IEEE Trans. Neural Networks Learn. Syst.6
2024 FLHetBench: Benchmarking Device and State Heterogeneity in Federated Learning
abstract
Federated learning (FL) is a powerful technology that enables collaborative training of machine learning models without sharing private data among clients. The fundamental challenge in FL lies in learning over extremely hetero-geneous data distributions, device capacities, and device state availabilities, all of which adversely impact performance and communication efficiency. While data hetero-geneity has been well-studied in the literature, this paper introduces FLHetBench, the first FL benchmark targeted toward understanding device and state heterogeneity. FL-HetBench comprises two new sampling methods to generate real-world device and state databases with varying het-erogeneity and new metrics for quantifying the success of FL methods under these real-world constraints. Using FL-HetBench, we conduct a comprehensive evaluation of existing methods and find that they struggle under these settings, which inspires us to propose BiasPrompt+, a new method employing staleness-aware aggregation and fast weights to tackle these new heterogeneity challenges. Experiments on various FL tasks and datasets validate the effectiveness of our BiasPrompt+ method and highlight the value of FLHet-Bench in fostering the development of more efficient and robust FL solutions under real-world device and state constraints.
Junyuan Zhang, Shuang Zeng, Miao Zhang 0030, Runxi Wang, Yuyin Zhou, Paul Pu Liang, Liangqiong Qu
CVPR2
2024 LUCIDA: Low-Dose Universal-Tissue CT Image Domain Adaptation for Medical Segmentation
Shuang Zeng, Zhaoheng Xie
MICCAI (8)4
2024 Low-Rank Mixture-of-Experts for Continual Medical Image Segmentation
Lei Zhu 0012, Hangzhou He, Shuang Zeng, Qiushi Ren, Yanye Lu
MICCAI (8)5
2024 Tackling Data Heterogeneity in Federated Learning via Loss Decomposition
Shuang Zeng, Pengxin Guo 0001, Yuyin Zhou, Liangqiong Qu
MICCAI (10)1
2024 Modeling and compensation of small-sample thermal error in precision machine tool spindles using spatial-temporal feature interaction fusion network
Xuesong Mei, Yuansheng Zhou, Jialan Liu, Shuang Zeng, Hongquan Gui, Jianqiang Zhou, Shengbin Weng
Adv. Eng. Informatics9
2024 Thermal error prediction of precision boring machine tools based on extreme gradient boosting algorithm-improved sailed fish optimizer-bi-directional ordered neurons-long short-term memory neural network model and physical-edge-cloud system
Jialan Liu, Hongquan Gui, Shuang Zeng, Fangqiong Luo
Eng. Appl. Artif. Intell.5
2024 SCA-YOLO: a new small object detection model for UAV images
Shuang Zeng, Wenzhu Yang, Yanyan Jiao, Lei Geng, Xinting Chen
Vis. Comput.1
2023 Coarse-to-Fine Entity Representations for Document-Level Relation Extraction
Damai Dai, Shuang Zeng, Baobao Chang, Zhifang Sui
NLPCC (2)3
2023 A multilayer human motion prediction perceptron by aggregating repetitive motion
Lei Geng, Wenzhu Yang, Yanyan Jiao, Shuang Zeng, Xinting Chen
Mach. Vis. Appl.4
2022 DISK: Domain-constrained Instance Sketch for Math Word Problem Generation
abstract
A math word problem (MWP) is a coherent narrative which reflects the underlying logic of math equations. Successful MWP generation can automate the writing of mathematics questions. Previous methods mainly generate MWP text based on inflexible pre-defined templates. In this paper, we propose a neural model for generating MWP text from math equations. Firstly, we incorporate a matching model conditioned on the domain knowledge to retrieve a MWP instance which is most consistent with the ground-truth, where the domain is a latent variable extracted with a domain summarizer. Secondly, by constructing a Quantity Cell Graph (QCG) from the retrieved MWP instance and reasoning over it, we improve the model’s comprehension of real-world scenarios and derive a domain-constrained instance sketch to guide the generation. Besides, the QCG also interacts with the equation encoder to enhance the alignment between math tokens (e.g., quantities and variables) and MWP text. Experiments and empirical analysis on educational MWP set show that our model achieves impressive performance in both automatic evaluation metrics and human evaluation metrics.
Tianyang Cao, Shuang Zeng, Xiaodan Xu, Mairgup Mansur, Baobao Chang
COLING2
2022 SCL-RAI: Span-based Contrastive Learning with Retrieval Augmented Inference for Unlabeled Entity Problem in NER
abstract
Unlabeled Entity Problem (UEP) in Named Entity Recognition (NER) datasets seriously hinders the improvement of NER performance. This paper proposes SCL-RAI to cope with this problem. Firstly, we decrease the distance of span representations with the same label while increasing it for different ones via span-based contrastive learning, which relieves the ambiguity among entities and improves the robustness of the model over unlabeled entities. Then we propose retrieval augmented inference to mitigate the decision boundary shifting problem. Our method significantly outperforms the previous SOTA method by 4.21% and 8.64% F1-score on two real-world datasets.
Shuzheng Si, Shuang Zeng, Jiaxing Lin, Baobao Chang
COLING2
2022 Type-enriched Hierarchical Contrastive Strategy for Fine-Grained Entity Typing
abstract
Fine-grained entity typing (FET) aims to deduce specific semantic types of the entity mentions in the text. Modern methods for FET mainly focus on learning what a certain type looks like. And few works directly model the type differences, that is, let models know the extent that which one type is different from others. To alleviate this problem, we propose a type-enriched hierarchical contrastive strategy for FET. Our method can directly model the differences between hierarchical types and improve the ability to distinguish multi-grained similar types. On the one hand, we embed type into entity contexts to make type information directly perceptible. On the other hand, we design a constrained contrastive strategy on the hierarchical structure to directly model the type differences, which can simultaneously perceive the distinguishability between types at different granularity. Experimental results on three benchmarks, BBN, OntoNotes, and FIGER show that our method achieves significant performance on FET by effectively modeling type differences.
Xinyu Zuo, Haijin Liang, Ning Jing, Shuang Zeng
COLING4
2022 Mining Clues from Incomplete Utterance: A Query-enhanced Network for Incomplete Utterance Rewriting
abstract
Incomplete utterance rewriting has recently raised wide attention.However, previous works do not consider the semantic structural information between incomplete utterance and rewritten utterance or model the semantic structure implicitly and insufficiently.To address this problem, we propose a QUEry-Enhanced Network (QUEEN).Firstly, our proposed query template explicitly brings guided semantic structural knowledge between the incomplete utterance and the rewritten utterance making model perceive where to refer back to or recover omitted tokens.Then, we adopt a fast and effective edit operation scoring network to model the relation between two tokens.Benefiting from extra information and the well-designed network, QUEEN achieves state-of-the-art performance on several public datasets.
Shuzheng Si, Shuang Zeng, Baobao Chang
NAACL-HLT2
2022 A Two-Stream AMR-enhanced Model for Document-level Event Argument Extraction
abstract
Runxin Xu, Peiyi Wang, Tianyu Liu, Shuang Zeng, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Runxin Xu, Peiyi Wang, Tianyu Liu 0001, Shuang Zeng, Baobao Chang, Zhifang Sui
NAACL-HLT4
2021 Generating Math Word Problems from Equations with Topic Consistency Maintaining and Commonsense Enforcement
Tianyang Cao, Shuang Zeng, Songge Zhao, Mairgup Mansur, Baobao Chang
ICANN (3)2
2020 Double Graph Based Reasoning for Document-level Relation Extraction
abstract
Document-level relation extraction aims to extract relations among entities within a document.Different from sentence-level relation extraction, it requires reasoning over multiple sentences across paragraphs.In this paper, we propose Graph Aggregation-and-Inference Network (GAIN), a method to recognize such relations for long paragraphs.GAIN constructs two graphs, a heterogeneous mentionlevel graph (MG) and an entity-level graph (EG).The former captures complex interaction among different mentions and the latter aggregates mentions underlying for the same entities.Based on the graphs we propose a novel path reasoning mechanism to infer relations between entities.Experiments on the public dataset, DocRED, show GAIN achieves a significant performance improvement (2.85 on F1) over the previous state-of-the-art.Our code is available at https://github.com/ PKUnlp-icler/GAIN.
Shuang Zeng, Runxin Xu, Baobao Chang, Lei Li 0005
EMNLP (1)1
2020 Evaluating Text Coherence at Sentence and Paragraph Levels
abstract
In this paper, to evaluate text coherence, we propose the paragraph ordering task as well as conducting sentence ordering. We collected four distinct corpora from different domains on which we investigate the adaptation of existing sentence ordering methods to a paragraph ordering task. We also compare the learnability and robustness of existing models by artificially creating mini datasets and noisy datasets respectively and verifying the efficiency of established models under these circumstances. Furthermore, we carry out human evaluation on the rearranged passages from two competitive models and confirm that WLCS-l is a better metric performing significantly higher correlations with human rating than τ , the most prevalent metric used before. Results from these evaluations show that except for certain extreme conditions, the recurrent graph neural network-based model is an optimal choice for coherence modeling.
Sennan Liu, Shuang Zeng, Sujian Li
LREC2