VLDB 2026 Research / reviewers in the wild / expert
Yicheng Feng
dblp:340/4016
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Compute and Network Orchestration for Disaggregated RLabstractDisaggregating the generation and training stages in RL is widely adopted to scale LLM post-training. There are two critical challenges here. First, the generation stage often becomes a bottleneck due to dynamic workload shifts and severe execution imbalances. Second, the decoupled stages result in diverse and dynamic network traffic patterns that strain the conventional static fabric. Xin Tan 0004, Yicheng Feng, Yu Zhou 0008, Yibo Zhu 0001, Hong Xu 0001 |
SIGCOMM | 2 |
| 2026 | A Co-Design Framework for Container Deployment in Mobile Edge Computing NetworksabstractWith the rapid advancement of mobile technologies, including self-driving cars and drones, the deployment of mobile software has become increasingly complex. In this context, virtualization plays a pivotal role by simplifying service deployment through containers and enabling container orchestration plat forms to efficiently manage an expanding number of container clusters. This is achieved by leveraging standardized interfaces and minimizing resource optimization overhead. However, the use of distributed servers in mobile edge clusters introduces several challenges, such as bandwidth limitations, network performance fluctuations, and resource constraints, which complicate deployment in these dynamic and resource-constrained environments. In this paper, we rethink the layer-based structure, a fundamental container design, and analyze the challenges and potential of real edge platform traces. Consequently, we propose BREAK, an acceleration middleware for efficient container deployment. With the primary insight of enhancing layer-reuse and deriving benefits from it, we develop a co-design approach centered on layer structure for efficient deployment, ensuring backward compatibility: (i) a container image refactoring solution that optimizes efficiency while preserving the stack-of-layers structure, (ii) distributed shared layer-stack caches, dynamically optimized for collaborative container deployment among mobile edge clusters, (iii) a customized Kubernetes (K8s) scheduler extending awareness of network performance, disk space, and container layer cache for container placement, and (iv) a tailored storage-driver of the standard container runtime for efficient layer extraction. Results indicate that BREAK accelerates the deployment process by up to 2.1× and reduces redundant image size by up to 3.11× compared to the state-of-the-art approach. Shihao Shen, Yicheng Feng, Xiaoxu Ren, Xiaofei Wang 0001, Qiao Xiang, Hong Xu 0001, Chenren Xu |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | VideoOrion: Tokenizing Object Dynamics in VideosabstractWe present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline, encoding them into a set of object tokens by aggregating spatial-temporal object features. Our method addresses the persistent challenge in Video-LLMs of efficiently compressing high-dimensional video data into semantic tokens that are comprehensible to LLMs. Compared to prior methods which resort to downsampling the original video or aggregating visual tokens using resamplers, leading to information loss and entangled semantics, VideoOrion not only offers a more natural and efficient way to derive compact, disentangled semantic representations but also enables explicit object modeling of video content with minimal computational cost. Moreover, the introduced object tokens naturally allow VideoOrion to accomplish video-based referring tasks. Experimental results show that VideoOrion can learn to make good use of the object tokens, and achieves competitive results on both general video question answering and video-based referring benchmarks. Yicheng Feng, Yijiang Li, Wanpeng Zhang 0002, Sipeng Zheng, Hao Luo 0011, Zihao Yue, Zongqing Lu 0002 |
ICCV | 1 |
| 2025 | Unified Multimodal Understanding via Byte-Pair Visual EncodingabstractMultimodal large language models (MLLMs) have made significant progress in vision-language understanding, yet effectively aligning different modalities remains a fundamental challenge. We present a framework that unifies multimodal understanding by applying byte-pair encoding to visual tokens. Unlike conventional approaches that rely on modality-specific encoders, our method directly incorporates structural information into visual tokens, mirroring successful tokenization strategies in text-only language models. We introduce a priority-guided encoding scheme that considers both frequency and spatial consistency, coupled with a multi-stage training procedure based on curriculum-driven data composition. These enhancements enable the transformer model to better capture cross-modal relationships and reason with visual information. Comprehensive experiments demonstrate improved performance across diverse vision-language tasks. By bridging the gap between visual and textual representations, our approach contributes to the advancement of more capable and efficient multimodal foundation models. Wanpeng Zhang 0002, Yicheng Feng, Hao Luo 0011, Yijiang Li, Zihao Yue, Sipeng Zheng, Zongqing Lu 0002 |
ICCV | 2 |
| 2025 | From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual ModalitiesabstractMultimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We introduce a novel image tokenizer that bridges this gap by applying the principle of Byte-Pair Encoding (BPE) to visual data. Unlike conventional approaches that rely on separate visual encoders, our method directly incorporates structural prior information into image tokens, mirroring the successful tokenization strategies used in text-only Large Language Models. This innovative approach enables Transformer models to more effectively learn and reason across modalities. Through theoretical analysis and extensive experiments, we demonstrate that our BPE Image Tokenizer significantly enhances MLLMs' multimodal understanding capabilities, even with limited training data. Leveraging this method, we develop Being-VL-0, a model that demonstrates superior performance across various benchmarks and shows promising scalability, potentially paving the way for more efficient and capable multimodal foundation models. For further details, visit our website https://github.com/BeingBeyond/Being-VL-0. Wanpeng Zhang 0002, Zilong Xie, Yicheng Feng, Yijiang Li, Xingrun Xing, Sipeng Zheng, Zongqing Lu 0002 |
ICLR | 3 |
| 2025 | OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and DataabstractRecent advances in large multimodal models have significantly advanced video comprehension, yet their performance remains limited in first-person scenarios. The interactive nature of egocentric videos is critical for applications like embodied intelligence, but introduces complex visual contexts that conventional models struggle to capture. To bridge this gap, we introduce OpenMMEgo with innovations across three dimensions: data, model, and training strategy. To provide rich spatiotemporal visual knowledge, we curate a large-scale, high-quality dataset named OME10M, comprising over 8.2M egocentric video QA pairs synthesized from Ego4D series. We also establish OMEBench, a comprehensive benchmark for rigorous egocentric understanding assessment. To alleviate the frequent viewpoint shifts inherent in egocentric videos, we implement semantic-aware visual token compression. Further, a curriculum learning strategy is complemented to foster stable learning across various data complexities. OpenMMEgo consistently improves the performance of LMMs on egocentric benchmarks without sacrificing general video understanding performance. Notably, Qwen2.5-VL tuned with OpenMMEgo substantially outperforms other models of the same size in egocentric video understanding. The data, weights and training code will be put at https://github.com/BeingBeyond/OpenMMEgo. Hao Luo 0011, Zihao Yue, Wanpeng Zhang 0002, Yicheng Feng, Sipeng Zheng, Deheng Ye, Zongqing Lu 0002 |
NeurIPS | 4 |
| 2025 | Levitation Performance of the High-Temperature Superconducting Scaled-Vehicle Under Magnetic Field Irregularity by Six Degrees of Freedom of Electromagnetic Force-Dynamics Coupling Model
Wuyang Lei, Can Peng, Peiyu Yin, Yicheng Feng, Zigang Deng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Learning Multi-Object Positional Relationships via Emergent CommunicationabstractThe study of emergent communication has been dedicated to interactive artificial intelligence. While existing work focuses on communication about single objects or complex image scenes, we argue that communicating relationships between multiple objects is important in more realistic tasks, but understudied. In this paper, we try to fill this gap and focus on emergent communication about positional relationships between two objects. We train agents in the referential game where observations contain two objects, and find that generalization is the major problem when the positional relationship is involved. The key factor affecting the generalization ability of the emergent language is the input variation between Speaker and Listener, which is realized by a random image generator in our work. Further, we find that the learned language can generalize well in a new multi-step MDP task where the positional relationship describes the goal, and performs better than raw-pixel images as well as pre-trained image features, verifying the strong generalization ability of discrete sequences. We also show that language transfer from the referential game performs better in the new task than learning language directly in this task, implying the potential benefits of pre-training in referential games. All in all, our experiments demonstrate the viability and merit of having agents learn to communicate positional relationships between multiple objects through emergent communication. Yicheng Feng, Boshi An, Zongqing Lu 0002 |
AAAI | 1 |
| 2024 | UniCode: Learning a Unified Codebook for Multimodal Large Language Models
Sipeng Zheng, Yicheng Feng, Zongqing Lu 0002 |
ECCV (8) | 3 |
| 2024 | Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open WorldsabstractRecent studies have presented compelling evidence that large language models (LLMs) can equip embodied agents with the self-driven capability to interact with the world, which marks an initial step toward versatile robotics. However, these efforts tend to overlook the visual richness of open worlds, rendering the entire interactive process akin to ``a blindfolded text-based game.'' Consequently, LLM-based agents frequently encounter challenges in intuitively comprehending their surroundings and producing responses that are easy to understand. In this paper, we propose Steve-Eye, an end-to-end trained large multimodal model to address this limitation. Steve-Eye integrates the LLM with a visual encoder to process visual-text inputs and generate multimodal feedback. We adopt a semi-automatic strategy to collect an extensive dataset comprising 850K open-world instruction pairs, enabling our model to encompass three essential functions for an agent: multimodal perception, foundational knowledge base, and skill prediction and planning. Lastly, we develop three open-world evaluation benchmarks and carry out experiments from a wide range of perspectives to validate our model's capability to strategically act and plan. The project’s website and code can be found at https://sites.google.com/view/steve-eye. Sipeng Zheng, Yicheng Feng, Zongqing Lu 0002 |
ICLR | 3 |
| 2024 | BREAK: A Holistic Approach for Efficient Container Deployment among Edge CloudsabstractContainer technology has revolutionized service deployment, offering streamlined processes and enabling container orchestration platforms to manage a growing number of container clusters. However, the deployment of containers in distributed edge clusters presents challenges due to their unique characteristics, such as bandwidth limitations and resource constraints. Existing approaches designed for cloud environments often fall short in addressing the specific requirements of edge computing. Additionally, very few edge-oriented solutions explore fundamental changes to the container design, resulting in difficulties achieving backward compatibility.In this paper, we reevaluate the fundamental layer-based structure of containers. We identify that the proliferation of redundant files and operations within image layers hinders efficient container deployment. Drawing upon the crucial insight of enhancing layer reuse and extracting benefits from it, we introduce BREAK, a holistic approach centered on layer structure throughout the entire container deployment pipeline, ensuring backward compatibility. BREAK refactors image layers and proposes an edge-oriented cache solution to enable ubiquitous and shared layers. Moreover, it addresses the complete deployment pipeline by introducing a customized scheduler and a tailored storage driver. Our results demonstrate that BREAK accelerates the deployment process by up to 2.1× and reduces redundant image size by up to 3.11× compared to state-of-the-art approaches. Yicheng Feng, Shihao Shen, Xiaofei Wang 0001, Qiao Xiang, Hong Xu 0001, Chenren Xu |
INFOCOM | 1 |
| 2024 | Tango: Harmonious Optimization for Mixed Services in Kubernetes-Based Edge CloudsabstractDeploying Latency-Critical (LC) services and Best-Effort (BE) services together is expected to improve resource utilization in edge clouds. However, co-locating LC and BE services on edge clouds presents unique challenges. Unlike cloud datacenters, edge clouds are heterogeneous, resource-constrained, and geographically distributed, leading to fiercer competition for resources and greater difficulty in balancing fluctuating co-located workloads. Due to the lack of consideration for the characteristics of edge environments, previous solutions designed for cloud datacenters are no longer applicable. To address these challenges, we introduceTango, a harmonious scheduling framework forKubernetes-based edge cloud systems with mixed services.Tangoincorporates novel components and mechanisms for elastic resource allocation on the edge, as well as two traffic scheduling algorithms that efficiently manage distributed edge resources.Tangofosters harmony not only by supporting compatible mixed services but also by offering collaborative solutions that complement each other. Based on a non-intrusive design forKubernetes,Tangofurther enhances it with automatic scaling and traffic scheduling capabilities. Compared to state-of-the-art approaches, experiments on large-scale hybrid edge clouds, driven by real workload traces, show thatTangoimproves system resource utilization by 36.9%, QoS-guarantee satisfaction rate by 11.3%, and throughput by 47.6%. Shihao Shen, Yicheng Feng, Mengwei Xu 0001, Yuanming Ren, Xiaofei Wang 0001, Victor C. M. Leung |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | Quicklayer: A Layer-Stack-Oriented Accelerating Middleware for Fast Deployment in Edge CloudsabstractContainers are gaining popularity in edge computing due to their standardization and low overhead. This trend has brought new technologies such as container engines and container orchestration platforms (COPs). However, fast and effective container deployment remains a challenge, especially at the edge. Prior work, which was designed for cloud datacenters, is no longer suitable for container deployment in edge clouds due to bandwidth limitations, fluctuating network performance, resource constraints, and geo-distributed organization. These edge features make rapid deployment on the edge difficult. Additionally, integrating with COPs is crucial for successful deployment. Yicheng Feng, Shihao Shen, Cheng Zhang 0019, Xiaofei Wang 0001 |
APNet | 1 |
| 2023 | Tango: Harmonious Management and Scheduling for Mixed Services Co-located among Distributed Edge-CloudsabstractCo-locating Latency-Critical (LC) and Best-Effort (BE) services in edge-clouds is expected to enhance resource utilization. However, this mixed deployment encounters unique challenges. Edge-clouds are heterogeneous, distributed, and resource-constrained, leading to intense competition for edge resources, making it challenging to balance fluctuating co-located workloads. Previous works in cloud datacenters are no longer applicable since they do not consider the unique nature of edges. Although very few works explicitly provide specific schemes for edge workload co-location, these solutions fail to address the major challenges simultaneously. Yicheng Feng, Shihao Shen, Mengwei Xu 0001, Yuanming Ren, Xiaofei Wang 0001, Victor C. M. Leung |
ICPP | 1 |
| 2023 | A Holistic QoS View of Crowdsourced Edge Cloud PlatformabstractEdge clouds have become a de-facto paradigm to deliver low and stable networks to delay-critical applications such as web services and AR/VR. A unique form of edge clouds is those crowdsourced from third parties, e.g., idle PCs or workstations. Such crowdsourced edge platforms can better sink computations closer to users, reduce the purchase cost, and eliminates the carbon generated during manufacturing. Yet, they also face the challenge of out-of-control hardware, e.g., a server dropping in/out anytime. In this paper, we perform the first-of-its-kind measurement of Quality of Service (QoS) for a large-scale crowdsourced edge platform, which covers over 10,000 edge servers, 100,000 users and 10,000,000 user requests. The measurement takes a holistic QoS view: (1) First, we look at how much hardware resources are provided by edge servers, how much time they are available for service deployment, and what are the major abnormal behaviors. (2) Second, we analyze the factors affecting service stability and quantify the resource utilization pattern of containerized services hosted on those edge servers. (3) Third, we investigate the spatial and temporal features of user requests handled by the platform. Many useful and somehow surprising findings are obtained through the above measurements. We also derive insightful implications that could help edge platforms and edge applications to better deliver their services to users. Shihao Shen, Yicheng Feng, Mengwei Xu 0001, Cheng Zhang 0007, Xiaofei Wang 0001, Victor C. M. Leung |
IWQoS | 2 |
| 2023 | A large-scale holistic measurement of crowdsourced edge cloud platform
Yicheng Feng, Shihao Shen, Mengwei Xu 0001, Cheng Zhang 0007, Xin Wang 0030, Xiaofei Wang 0001, Victor C. M. Leung |
World Wide Web (WWW) | 1 |