Siming Sun

dblp:80/10440 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-3037-118XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Markovian Linguistic-Temporal Bridge: Unlocking the Potential of LLMs for Time Series Forecasting
abstract
Adapting pretrained Large Language Models (LLMs) for time series forecasting primarily relies on token-level linguistic-temporal alignment, leading to the stacking of logically disjointed tokens as input.While empirically effective, these methods overlook a fundamental capability of LLMs: modeling linguistic logic and structure, rather than merely processing token features.To address this limitation, we propose the Markovian-Guided Structure-Aware Alignment (MGSAA).Our core contribution is a framework that transcends pointwise feature matching to achieve global structural isomorphism between the linguistic and temporal domains.Specifically, MGSAA distills latent evolutionary patterns of language within LLMs into a Markovian state transition graph, which is transferred as a structural prior to the time series domain.Under this prior, time series patches are decoded into latent states and then aligned via state-constrained cross-attention.Ultimately, MGSAA generates a token sequence topologically isomorphic to the LLM's inherent mental structure, reactivating its reasoning capabilities for forecasting.Comprehensive evaluations across multiple benchmarks demonstrate that MGSAA achieves state-of-the-art performance, providing an innovative solution for cross-modal alignment in LLM for time series forecasting.
Siming Sun, Kai Zhang 0079, Xuejun Jiang, Wenchao Meng, Qinmin Yang
ACL (1)1
2025 Code Design in Almost Lossless Onion Peeling Data Compression
abstract
Onion peeling codes address a distributed data compression scenario related to Slepian-Wolf (SW) data compression. In a 2-user onion peeling code, as in a 2-user SW code, two dependent sources are compressed independently; while the SW code decodes both sources jointly using both source descriptions, the onion peeling code reconstructs the first source using only the first source description and then uses the first source reconstruction and the second source description to reconstruct the second source. The authors' prior results show that first-stage compressors that are almost identical in first-source reconstruction reliability and efficiency can exhibit extremely different best-case performance for the second source. This paper proposes the use of low density parity check (LDPC) source coding in the first stage of an onion peeling code and shows that, with high probability, the random LDPC code design used in the evaluation of this approach generates a first-stage code that performs well both in compressing the first source and in assisting the compressor of the second source. The method works universally for any conditional distribution on the second source given the first, meaning that one does not have to know the conditional distribution of the second source given the first to design a first-stage code with good performance on the second source.
Siming Sun, Michelle Effros
DCC1
2025 Meta-Learning-Based Safety-Critical Control in Multi-Obstacles Environments
abstract
Autonomous robots operating in diverse scenarios are expected to safely and efficiently adapt to new, unknown, and cluttered environments. In this paper, we introduce a real-time goal-seeking and exploration framework incorporating novel meta-signed distance functions (MetaSDFs) and metabuffer robust control barrier functions (Meta-BRCBFs). To adapt to environmental changes in real time, we employ Bayesian meta-learning to construct MetaSDFs. Deep neural network weights are initially trained offline, followed by efficient online adaptation at the last Bayesian layer, allowing for online updates at linear time complexity. Each MetaSDF is individually trained for its corresponding obstacle class, enhancing online distance estimation accuracy. Subsequently, buffer zones are constructed around the MetaSDFs to establish corresponding Meta-BRCBFs. These Meta-BRCBFs are activated only when the robot enters these zones, substantially reducing the number of CBFs required. Outside these specified buffer zones, the robot remains ingoal-seekingmode, focusing on task completion. After entering a buffer zone, it transitions toexplorationmode, prioritizing safety and exploring safe pathways, effectively balancing task execution with environmental adaptability. We demonstrate that, under this framework, the system achieves both safety and asymptotic stabilization. Extensive simulations and experiments are conducted to demonstrate our framework’s effectiveness in both simulated scenarios and real-world environments. These tests confirm our framework’s real-time capabilities and safety assurances in dynamic settings where state-of-the-art methods fail. The video is available at: https://www.youtube.com/watch?v=C6eshldAMxA.
Yu Zhang 0182, Long Wen 0003, Yuhong Huang, Siming Sun, Zhenshan Bing, Wei He 0001, Alois C. Knoll
IEEE Trans Autom. Sci. Eng.4
2024 Almost Lossless Onion Peeling Data Compression
abstract
This work considers almost lossless onion peeling data compression. In onion peeling data compression, as in Slepian-Wolf data compression, multiple transmitters independently encode their respective sources and transmit their descriptions to a shared decoder. Onion peeling codes differ from Slepian-Wolf codes in that the onion peeling decoder must sequentially decode the individual sources rather than making a single joint decoding decision. This work considers an almost lossless two-stage code. The main result shows that when the first source is coded by a near-optimal code for a given first-stage error constraint, the conditional entropy rate of the second source, given this imperfect reconstruction, can vary within a gap that does not vanish when the blocklength grows without bound. It is also shown that the lower bound for the second-stage conditional entropy rate can be achieved using a classic random binning code in the first stage.
Siming Sun, Michelle Effros
DCC1
2024 Asynchronous Random Access Data Compression
abstract
This work introduces a framework for an asynchronous random access source code (ARASC) and bounds the achievable performance under this framework. Like prior multiple access (or Slepian-Wolf) source codes, the ARASC enables multiple transmitters to efficiently, reliably, and independently describe dependent sources to a common receiver. As in prior "random access" codes, the number of active encoders is unknown a priori to both the transmitters and the receiver and single-bit stop-feedback from the receivers to the transmitters enables variable-rate coding. Unlike prior works, the proposed system eliminates all forms of block synchronization. The main result is a two-transmitter achievability bound demonstrating the achievability of a first-order average rate across blocks equal to the weighted average of the point-to-point source coding rate and multiple access achievable sum rate. The weights observed approach the fractions of time that separate and simultaneous observations are encoded. The result’s second order term bounds the speed at which the average rate approaches this weighted average.
Siming Sun, Michelle Effros
DCC1
2022 Source Coding with Unreliable Side Information in the Finite Blocklength Regime
abstract
This paper studies a special case of the problem of source coding with side information. A single transmitter describes a source to a receiver that has access to a side information observation that is unavailable at the transmitter. While the source and true side information sequences are dependent, stationary, memoryless random processes, the side information observation at the decoder is unreliable, which here means that it may or may not equal the intended side information and therefore may or may not be useful for decoding the source description. The probability of side information observation failure, caused, for example, by a faulty sensor or source decoding error, is non-vanishing but is bounded by a fixed constant independent of the blocklength. This paper proposes a coding system that uses unreliable side information to get efficient source representation subject to a fixed error probability bound. Results include achievability and converse bounds under two different models of the joint distribution of the source, the intended side information, and the side information observation.
Siming Sun, Michelle Effros
ISIT1
2011 Exploring the corporate ecosystem with a semi-supervised entity graph
abstract
Investment decisions in the financial markets require careful analysis of information available from multiple data sources. In this paper, we present Atlas, a novel entity-based information analysis and content aggregation platform that uses heterogeneous data sources to construct and maintain the "ecosystem" around tangible and logical entities such as organizations, products, industries, geographies, commodities and macroeconomic indicators. Entities are represented as vertices in a directed graph, and edges are generated using entity co-occurrences in unstructured documents and supervised information from structured data sources. Significance scores for the edges are computed using a method that combines supervised, unsupervised and temporal factors into a single score. Important entity attributes from the structured content and the entity neighborhood in the graph are automatically summarized as the entity "fingerprint". A highly interactive user interface provides exploratory access to the graph and supports common business use cases. We present results of experiments performed on five years of news and broker research data, and show that Atlas is able to accurately identify important and interesting connections in real-world entities. We also demonstrate that Atlas entity fingerprints are particularly useful in entity similarity queries, with a quality that rivals existing human maintained databases.
Hassan H. Malik, Ian MacGillivray, Måns Olof-Ors, Siming Sun, Shailesh Saroha
CIKM4