Vadim Sheinin

dblp:78/738 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0003-0278-2483ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Agentic Solutions for IT Financial Operations
abstract
The dynamic nature of cloud spending and pricing structures pose challenges for practitioners in IT Financial Operations (FinOps). Recent advances in agentic systems enables them to instead rely on agents for complex FinOps tasks such as drawing insights from their data through natural language queries. In this work, we present an IT FinOps Data Insights Agent, that implements “chat with your data” approach to support practitioners in their daily tasks. Our agent achieves up to 90% accuracy across ITBench FinOps scenarios.
Bekir O. Turkkan, Pavankumar Murali, Chandrasekhar Narayanaswami 0001, Vadim Sheinin
AAAI4
2024 A Hybrid Cognitive Contract Application for Identifying Accounting Risks in Contractual Language
Ngoc Phuoc An Vo, Martin Linhart, Fruzsina Strbik, Istvan Koska, Petros Zerfos, Vadim Sheinin, Jeff Dakin, Milton Laverde
NLDB (2)6
2022 Natural Language Interface for Process Mining Queries in Healthcare
abstract
Recently, the needs of data required for data analysis are becoming more diversified, and research on data extraction and analysis methods has been continuously made in order to effectively respond to various needs. Process mining is a solution that analyzes various system logs built by companies or healthcare institutions so that they can be used for process improvement. From the process model extracted from the system logs, it is possible not only to grasp the exact flow of the current business process, but also to acquire additional information such as repetitive execution of activities in the process where the bottleneck occurs in the business process flow. The manufacturing industry has made great efforts to improve the process management, and as many companies are paying attention to big data these days, various data-related technologies are emerging in the healthcare industry as well to properly provide patients with the care needed. Process mining tools allow users to pull data by programming in a process mining query language using the APIs provided with the process mining tool, or by manually creating reusable analytical documents using user friendly tool. However, these tasks require the users to be familiar with the query language APIs and understand the data model and its relationships with respect to creating analytical documents. This paper proposes a methodology that allows users to easily extract desired data through natural language interface, which relieves nonprofessional users of the burden of programming in a process mining query language. The process mining query engine with natural language interface presented in this paper consists of four major components. Among them, the natural language processing pipeline that not only extracts intermediate representation of entities used when constructing a process mining query language report from natural language queries, but also effectively extracts a query hint from the context of natural language query. The query hint is used to select a process-specific function from the library that fits the context of the user query while transforming a natural language query into a process mining query report. The method proposed in this study has the advantage of being able to roughly grasp the process state for the user just by entering a query in natural language. The proposed system provides users with four query process options. That is, the user 1) retrieves intermediate representation of entities and query hints from the NLP pipeline, 2) retrieves the process mining query language from the query language generator, 3) submits the query language to the process mining engine and execute the query, 4) retrieves description of intermediate representation of entity and query hints in natural language to confirm that the query is processed correctly. The contents proposed in this paper were constructed and executed, and the query reports in process mining query language programmatically generated by the proposed query engine were also executed in a process mining engine and the query results were verified.
Hangu Yeo, Elahe Khorasani, Vadim Sheinin, Irene Manotas, Ngoc Phuoc An Vo, Octavian Popescu, Petros Zerfos
IEEE Big Data3
2022 Addressing Limitations of Encoder-Decoder Based Approach to Text-to-SQL
abstract
Most attempts on Text-to-SQL task using encoder-decoder approach show a big problem of dramatic decline in performance for new databases. For the popular Spider dataset, despite models achieving 70% accuracy on its development or test sets, the same models show a huge decline below 20% accuracy for unseen databases. The root causes for this problem are complex and they cannot be easily fixed by adding more manually created training. In this paper we address the problem and propose a solution that is a hybrid system using automated training-data augmentation technique. Our system consists of a rule-based and a deep learning components that interact to understand crucial information in a given query and produce correct SQL as a result. It achieves double-digit percentage improvement for databases that are not part of the Spider corpus.
Octavian Popescu, Irene Manotas, Ngoc Phuoc An Vo, Hangu Yeo, Elahe Khorashani, Vadim Sheinin
COLING6
2021 Programmatic Database Language Generation for Big Data Applications
abstract
Database management systems offer an efficient way of managing huge amount of data such as financial and healthcare data and he data retrieval from databases requires knowledge of Structured Query Language (SQL). In this paper, an Automatic SQL Generation System is proposed to help users who are inexperienced in querying database with SQL. The proposed SQL generation system reads formatted data items in the query report from the user and converts the data items into SQL statements programmatically with the help of a data model that is pulled from a database. The SQL generation system can handle simple queries composed of a query block with a SELECT statement as well as complex queries composed of multiple query blocks containing multiple SELECT statements. The proposed system is integrated with an NLIDB (Natural Language Interface for Database) system to translate data items (or tokens) extracted from queries in natural languages into SQL query language, and the system is also integrated and adapted with various types of databases and use cases that include financial and healthcare use cases. The experiment results show that the proposed system correctly handles user queries in natural language just like any other neural model based system and more importantly, the proposed SQL generation engine generates SQL queries without syntactic problems with various databases for all queries.
Hangu Yeo, Elahe Khorasani, Vadim Sheinin, Ngoc Phuoc An Vo, Octavian Popescu, Petros Zerfos
IEEE BigData3
2020 Identifying Motion Entities in Natural Language and A Case Study for Named Entity Recognition
abstract
Motion recognition is one of the basic cognitive capabilities of many life forms, however, detecting and understanding motion in text is not a trivial task.In addition, identifying motion entities in natural language is not only challenging but also beneficial for a better natural language understanding.In this paper, we present a Motion Entity Tagging (MET) model to identify entities in motion in a text using the Literal-Motion-in-Text (LiMiT) dataset for training and evaluating the model.Then we propose a new method to split clauses and phrases from complex and long motion sentences to improve the performance of our MET model.We also present results showing that motion features, in particular, entity in motion benefits the Named-Entity Recognition (NER) task.Finally, we present an analysis for the special co-occurrence relation between the person category in NER and animate entities in motion, which significantly improves the classification performance for the person category in NER.
Ngoc Phuoc An Vo, Irene Manotas, Vadim Sheinin, Octavian Popescu
COLING3
2019 Tackling Complex Queries to Relational Databases
Octavian Popescu, Ngoc Phuoc An Vo, Vadim Sheinin, Elahe Khorashani, Hangu Yeo
ACIIDS (1)3
2019 A Natural Language Interface Supporting Complex Logic Questions for Relational Databases
Ngoc Phuoc An Vo, Octavian Popescu, Vadim Sheinin, Elahe Khorasani, Hangu Yeo
NLDB3
2018 Unsupervised Threshold Autoencoder to Analyze and Understand Sentence Elements
abstract
Analysis of legal and contract documents often requires both the discovery of document structure, as well as the accurate identification of important elements such as party (buyer, supplier), nature (obligation, right) and category (warranties, delivery, etc.). Hence, exploring novel features that lead to better element classification accuracy as well as better document structure discovery is particularly important. In this paper, we develop and present novel unsupervised learning techniques to analyze a large scale corpus of contract documents with the goal of learning and deriving new features to enhance classification accuracy over the elements of interest, and to extract relevant features leading to meaningful clusterings over contract structures. Particularly, we propose a novel t-threshold autoencoder neural network that flexibly controls the number of active neurons in response to sentences of different lengths at the network's input. Such an adaptive sparseness threshold enforces competition and specialization among encoding neurons and hence results in better features learning. We also present an extension of the convolutional neural network classifier that allows for the incorporation of these novel augmented features and show that higher classification accuracies over various classes of contract elements can be achieved. We further present a practical pipeline of deriving features from contract documents along with a clustering solution based on the K-means algorithm that leads to the separation among different types of sentences in the contract documents. We empirically demonstrate the performance of our developed techniques on a novel data corpus of Software Procurement contracts.
Xuan-Hong Dang, Raji Akella, Somaieh Bahrami, Vadim Sheinin, Petros Zerfos
IEEE BigData4
2018 SQL-to-Text Generation with Graph-to-Sequence Model
abstract
Previous work approaches the SQL-to-text generation task using vanilla Seq2Seq models, which may not fully capture the inherent graph-structured information in SQL query.In this paper, we first introduce a strategy to represent the SQL query as a directed graph and then employ a graph-to-sequence model to encode the global structure information into node embeddings.This model can effectively learn the correlation between the SQL query pattern and its interpretation.Experimental results on the WikiSQL dataset and Stackoverflow dataset show that our model significantly outperforms the Seq2Seq and Tree2Seq baselines, achieving the state-of-the-art performance. * Work done when the author
Kun Xu 0005, Lingfei Wu 0001, Zhiguo Wang 0006, Yansong Feng 0002, Vadim Sheinin
EMNLP5
2018 Exploiting Rich Syntactic Information for Semantic Parsing with Graph-to-Sequence Model
abstract
Existing neural semantic parsers mainly utilize a sequence encoder, i.e., a sequential LSTM, to extract word order features while neglecting other valuable syntactic information such as dependency or constituent trees.In this paper, we first propose to use the syntactic graph to represent three types of syntactic information, i.e., word order, dependency and constituency features; then employ a graph-tosequence model to encode the syntactic graph and decode a logical form.Experimental results on benchmark datasets show that our model is comparable to the state-of-the-art on Jobs640, ATIS, and Geo880.Experimental results on adversarial examples demonstrate the robustness of the model is also improved by encoding more syntactic information.
Kun Xu 0005, Lingfei Wu 0001, Zhiguo Wang 0006, Mo Yu, Vadim Sheinin
EMNLP6
2018 A Large Resource of Patterns for Verbal Paraphrases
Octavian Popescu, Ngoc Phuoc An Vo, Vadim Sheinin
LREC3
2018 QUEST: A Natural Language Interface to Relational Databases
Vadim Sheinin, Elahe Khorasani, Hangu Yeo, Ngoc Phuoc An Vo, Octavian Popescu
LREC1
2017 Ranking the importance of ontology concepts using document summarization techniques
abstract
Automated Ontology Learning systems are nowadays practical and used in a variety of domains. By using these systems, subject matter experts (SMEs) and ontology designers can readily construct very large ontologies consisting of tens of thousands of concepts and their relations based on a corpus. However, ontologies of this size make it extremely challenging for such SMEs to understand and further tune these ontologies. Prior studies have proposed techniques for concept ranking based solely on the analysis of the structure of the ontology graphs. In this paper, we propose a novel approach, which further exploits a word-level summarization technique applied to the source documents used to generate the ontology. Using the document summarization technique, we devise features that measure concept importance based on source documents where concepts are extracted. We demonstrate the effectiveness of our approach by comparing with existing ranking methods and by devising a scalable evaluation process inspired from the document retrieval domain.
Petros Zerfos, Vadim Sheinin, Nancy Greco
IEEE BigData3
2015 SDFS: Secure distributed file system for data-at-rest security for Hadoop-as-a-service
abstract
Cloud service providers are offering the popular Hadoop analytics platform following an "as-a-service" model, i.e. clusters of machines in their cloud infrastructures pre-configured with Hadoop software. Such offerings lower the cost and complexity of deploying a comparable system on-premises, however security considerations and in particular data confidentiality hamper wider adoption of such services by enterprises that handle data of sensitive nature. In this paper, we describe our efforts in providing security for data-at-rest (i.e. data that is stored) when Hadoop is offered as a cloud service. We analyze the requirements and architecture for such service and further describe a new distributed file system that we developed for Hadoop called SDFS, towards supporting this premise. We analyze parameter tuning for SDFS and through experiments on a real test-bed we evaluate its performance. We further present simulation results that explore the parameter space and can guide tuning.
Petros Zerfos, Hangu Yeo, Brent Paulovicks, Vadim Sheinin
IEEE BigData4
2011 Parallel Implementation of External Sort and Join Operations on a Multi-core Network-Optimized System on a Chip
Elahe Khorasani, Brent Paulovicks, Vadim Sheinin, Hangu Yeo
ICA3PP (1)3
2011 High performance computing of line of sight viewshed
abstract
In this paper we present our recent research and development work for multicore computing of Line of Sight (LoS) on the Cell Broadband Engine (CBE) processors. LoS can be found in many applications where real-time high performance computing is required. We will describe an efficient LoS multi-core parallel computing algorithm, including the data partition and computation load allocation strategies to fully utilize the CBE's computational resources for efficient LoS viewshed parallel computing. In addition, we will also illustrate a successive fast transpose algorithm to prepare the input data for efficient Single-Instruction-Multiple-Data (SIMD) operations. Furthermore, we describe the data input and output (I/O) management scheme to reduce the (I/O) latency in Direct-Memory-Access (DMA) data fetching and storing operations. The performance evaluation of our LoS viewshed computing scheme over an area of interest (AOI) with more than 4.19 million points has shown that our parallel computing algorithm on CBE takes less than 25.5 ms, which is several orders of magnitude faster than the available commercial systems.
Ligang Lu, Brent Paulovicks, Michael Perrone, Vadim Sheinin
ICME4
2010 Cell blade based H.264 video encoding engine for large scale video surveillance applications
abstract
Video surveillance has become one of the most important tools for public safety and security. In this paper, we present our new work on developing efficient parallel computing algorithms and schemes for implementing H.264 video encoder engine on Cell Broadband Engine (CBE) blade for high performance large scale video surveillance applications. Extending our previous work on H.264 video encoding on CBE, we have developed new parallel computing schemes in computational load partition, dynamic task scheduling, motion estimation, and mode selection. We partition the intensive H.264 video encoding computation load into four major functional modules, namely, the pre-processing module, the motion estimation module, the mode selection and transform/quantization module, and the Context Adaptive Binary Arithmetic Coding (CABAC) module. The task scheduler dynamically assigns a waiting computing task to a SPE as soon as it becomes available. Our new implementation has achieved more than 5X performance improvement to encode 32 standard-definition (SD 720x480 pixel resolution) H.264 video streams simultaneously at 30 frames per second with one Cell Blade that consists of 16 Synergistic Processor Elements (SPEs) and two control Power Processor Elements (PPEs) or 448 SD channels of H.264 video streams on a single chassis Cell Blade Center with 14 Cell Blades.
Ligang Lu, Brent Paulovicks, Vadim Sheinin, Michael Perrone
VCIP3
2008 Low-rate hybrid Wyner-Ziv coding of Laplace-Markov source using uniform scalar quantization
abstract
Hybrid Wyner-Ziv coders which employ a combination of Wyner-Ziv coding and differential pulse code modulation (DPCM) encoding have recently gained popularity for applications such as video coding. In this paper we analyze the low-rate operational rate distortion performance of Wyner-Ziv coding using uniform scalar quantization, in the context of such hybrid coders. Motivated by video we consider the compression of a first-order Laplace-Markov source, and derive approximate analytical rate and distortion expressions which are accurate at low rates. We utilize the derived analytical expressions to address the problem of determining the optimal quantization interval ratio of the Wyner-Ziv and DPCM scalar quantizers, for a range of rates.
Vadim Sheinin, Ashish Jagmohan, Dake He
ICASSP1
2008 On the Operational Rate-Distortion Performance of Uniform Scalar Quantization-Based Wyner-Ziv Coding of Laplace-Markov Sources
abstract
Wyner-Ziv (WZ) coding has recently been proposed as a low encoding complexity alternative to traditional DPCM coding for compression of sources with memory, in particular, in applications like multimedia compression. The viability of this alternative approach clearly depends on the compression performance of WZ coding compared to that of DPCM coding. In an attempt to understand the performance gap between WZ coding and DPCM coding, this paper studies the operational rate-distortion performance of WZ coding, using uniform scalar quantization followed by perfect Slepian-Wolf coding, for compression of a Laplace-Markov (LM) source. It is shown that at low rates or for weakly correlated LM sources, WZ coding is indeed a competitive alternative to DPCM coding. However, at high rates the performance gap becomes non-negligible for strongly correlated LM sources. In order to reduce the gap at high rates, a hybrid approach that combines DPCM coding and WZ coding is further investigated. It is shown that the hybrid approach is indeed competitive to DPCM coding at all rates even for strongly correlated LM sources.
Vadim Sheinin, Ashish Jagmohan, Dake He
IEEE Trans. Multim.1
2007 On the Performance of Uniform Threshold Quantization for a sum of Independent Memoryless Laplacian Sources
abstract
The performance of uniform threshold quantization subject to an entropy constraint is studied for a sum of two and three independent zero-mean memoryless Laplacian sources. Both symmetric and asymmetric quantizers are considered, and approximate parametric expressions for the operational rate-distortion function R(D) are obtained for all rates. In particular, the low rate regime (rates below 1 bit per sample) is considered and simpler expressions for R(D) are derived. It is envisioned that these expressions will facilitate rate estimation in practical Wyner-Ziv coding with low complexity encoder.
Vadim Sheinin, Dake He
ICASSP (3)1
2007 Uniform Scalar Quantization Based Wyner-Ziv Coding of Laplace-Markov Source
abstract
Wyner-Ziv coding has recently emerged as an alternative to conventional DPCM coding for compression of sources with memory, particularly in video compression. This paper studies the operational rate-distortion performance of Wyner-Ziv coding, using uniform scalar quantization followed by perfect Slepian-Wolf coding, for compression of a Laplace-Markov source. The performance gap of this technique relative to DPCM coding is characterized through derived rate-distortion expressions and numerical simulations.
Vadim Sheinin, Ashish Jagmohan, David He
ICASSP (1)1
2007 Motion Estimation with Similarity Constraint and its Application to Distributed Video Coding
abstract
In this paper we present a new motion estimation scheme that minimizes the objective distance function with a constraint of similarity measure to exploit the motion correlation among adjacent pixel blocks with similar statistics features. We formulate this correlation as a similarity measure on the motion vectors between the current pixel block and its neighboring blocks weighted by the corresponding statistical similarity. We then use this similarity measure as a constraint in the objective distance function to reduce the noise effects and improve performance by effectively trading off the difference in the pixel values with the smoothness in the adjacent motion vectors. Thus, in motion estimation, our new scheme not only minimizes the pixel differences but tries to preserve the motion smoothness among statistically similar neighbors. We applied this new motion estimation scheme to a distributed video coding system for side information generation for Wyner-Ziv decoding and compared its performance to the scheme without similarity constraint. The results have shown that our motion estimation scheme can achieve significant gains in the fidelity of the side information and the decoded Wyner-Ziv frames over the scheme without the similarity constraint.
Ligang Lu, Vadim Sheinin
ICME2
2007 Accelerating Mutual-Information-Based Linear Registration on the Cell Broadband Engine Processor
abstract
Emerging multi-core processors are able to accelerate medical imaging applications by exploiting the parallelism available in their algorithms. We have implemented a mutual-information-based 3D linear registration algorithm on the Cell Broadband Enginetrade processor. By exploiting the highly parallel architecture and its high memory bandwidth, our implementation with two CBE processors can register a pair of 256x256x30 3D images in one second. This implementation is significantly faster than a conventional one on a traditional microprocessor or even faster than a previously reported custom-hardware implementation. In addition to parallelizing the code for multiple cores and organizing the data structure for reducing the amount of the memory traffic, it is also critical to optimize the code for the SIMD pipeline structure. We note that code optimization for the SIMD pipeline alone results in a 4.2x-8.7x acceleration for the computation of small kernels. Further, SIMD optimization alone results in a 4.5x end-end application speedup.
Moriyoshi Ohara, Hangu Yeo, Frank Savino, Giridharan Iyengar, Leiguang Gong, Hiroshi Inoue, Hideaki Komatsu, Vadim Sheinin, Shahrokh Daijavad
ICME8
2007 On A Partial Ordering Relation Derived from Redundancy of Slepian-Wolf Coding
abstract
Let (X, Y) denote a pair of finite-valued random variables. In this paper we use two examples to show an inherent partial ordering relation among the set {Py\x: H(X\Y) = a} where {Py\x : H(X\Y) = a} denotes the channel from X to Y, and 0 lesplusmn les H(X) is a constant. Specifically, we consider the following cases: the channel from X to Y is either a binary symmetric channel (BSC) or a binary erasure channel (BEC). In each case, we characterize the redundancy of Slepian-Wolf coding of X with decoder only side information Y. It is thus revealed that for any binary X and 0 < a < H(X), under the condition that H(X\Y) = a the redundancy of the BSC case is strictly larger than that of the BEC case for a range of decoding error probabilities. Interestingly, our results also reveal that the redundancy of variable-rate Slepian-Wolf coding is generally better than that of fixed-rate Slepian-Wolf coding.
Dake He, Ashish Jagmohan, Vadim Sheinin
ISIT3
2006 Uniform Threshold Scalar Quantizer Performance in Wyner-Ziv Coding With Memoryless, Additive Laplacian Correlation Channel
abstract
The performance of a uniform-threshold scalar quantizer in Wyner-Ziv coding is investigated in this paper. To derive analytical expressions we assume the abstract correlation channel from the side information to the source to be encoded is memoryless, additive Laplacian. Furthermore, in order to focus our attention on the performance of the quantizer, the Wyner-Ziv coding scheme is assumed to encode the quantizer output by using perfect Slepian-Wolf coding. Analytical expressions for the operational rate-distortion function are obtained for this case. By evaluating these analytical expressions, we show that scalar quantization with a mid-tread uniform threshold quantizer, followed by perfect Slepian Wolf coding achieves performance which is close to the theoretical Wyner-Ziv rate-distortion bound at low rates
Vadim Sheinin, Ashish Jagmohan, Dake He
ICASSP (4)1
2006 Low Rate Uniform Scalar Quantization of Memoryless Gaussian Sources
abstract
The low-rate (<;1 bits per sample) operational rate-distortion performance of uniform scalar quantizers for the memoryless Gaussian source is studied. Approximate analytical expressions for the operational rate-distortion function are derived, and the accuracy of the derived function is verified through simulation. It is shown that in the zero-rate limit the derived operational rate-distortion function is first-order optimal with respect to the Shannon lower bound. The derived function is used to study the performance of uniform scalar quantizers for the Gaussian Wyner-Ziv problem. Lastly, the derived low-rate rate-distortion function is used to provide improved low-rate bit allocation for jointly Gaussian vectors.
Vadim Sheinin, Ashish Jagmohan
ICIP1
2006 Video Analysis and Compression on the STI Cell Broadband Engine Processor
abstract
With increased concern for physical security, video surveillance is becoming an important business area. Similar camera-based system can also be used in such diverse applications as retail-store shopper motion analysis and casino behavioral policy monitoring. There are two aspects of video surveillance that require significant computing power: image analysis for detecting objects, and video compression for digital storage. The new STI CELL broadband engine (CBE) processor is an appealing platform for such applications because it incorporates 8 separate high-speed processing cores with an aggregate performance of 256Gflops. Moreover, this chip is the heart of the new Sony Playstation 3 and can be expected to be relatively inexpensive due to the high volume of production. In this paper we show how object detection and compression can be implemented on the CBE, discuss the difficulties encountered in porting the code, and provide performance results demonstrating significant speed-up
Lurng-Kuo Liu, Sreeni Kesavarapu, Jonathan H. Connell, Ashish Jagmohan, Lark-hoon Leem, Brent Paulovicks, Vadim Sheinin, Lijung Tang, Hangu Yeo
ICME7
2004 Rate and decoding power constrained video coding scheme for mobile multimedia players
abstract
Power dissipation, data rate, and processing time are crucial constraints for embedded multimedia devices. We present a video coding solution given the constraints of bit rate and decoding power dissipation. Coupled with a small dedicated video internal memory at a decoder, the coding scheme is designed for the best operational rate-distortion and decoding power cost trade off. It can substantially decrease the data traffic to the external memory at decoder. A decrease in data traffic to the external memory at decoder will result in faster real-time processing and power savings. The encoder, given the prior knowledge of the decoder's dedicated video internal memory management scheme, optimally regulates its choice of motion compensated predictors to reduce the decoder's external memory accesses. This video coding scheme can be used in any standard or proprietary encoder to generate a compliant output stream decodable by a standard general purpose processor-based or dedicated hardware-based decoder with power restriction. Simulation results show that with a relatively small dedicated internal memory, our scheme may reduce the power dissipation significantly while keeping the comparable picture quality.
Ligang Lu, Vadim Sheinin
ICIP2
2004 Video coding for decoding power-constrained embedded devices
abstract
Low power dissipation and fast processing time are crucial requirements for embedded multimedia devices. This paper presents a technique in video coding to decrease the power consumption at a standard video decoder. Coupled with a small dedicated video internal memory cache on a decoder, the technique can substantially decrease the amount of data traffic to the external memory at the decoder. A decrease in data traffic to the external memory at decoder will result in multiple benefits: faster real-time processing and power savings. The encoder, given prior knowledge of the decoder’s dedicated video internal memory cache management scheme, regulates its choice of motion compensated predictors to reduce the decoder’s external memory accesses. This technique can be used in any standard or proprietary encoder scheme to generate a compliant output bit stream decodable by standard CPU-based and dedicated hardware-based decoders for power savings with the best quality-power cost trade off. Our simulation results show that with a relatively small amount of dedicated video internal memory cache, the technique may decrease the traffic between CPU and external memory over 50%.
Ligang Lu, Vadim Sheinin
VCIP2
2003 Real-time MPEG video coding with information look-ahead
abstract
There have been increasing market demands to develop more effective schemes for applying MPEG video coding standard to meet the challenges from the emerging applications. We present a real-time MPEG video coding system with information look-ahead for constant bit rate (CBR) applications, such as Video-on-Demand (VoD) over ADSL. This scheme employs two MPEG encoders. Both encoders operate at the same CBR. The second encoder has a buffer to delay the input by an amount of time relative to the first encoder to create a look-ahead window. In encoding, the first encoder collects the information of statistics and rate-quality characteristics. An on-line information processor then uses the collected information to derive the best coding strategy for the second encoder to encode the incoming frames in the look-ahead window. The second encoder uses the encoding parameters from the processor as the coding guide to execute the coding strategy and generate the final bitstream. Results on an IBM MPEG-2 encoder testing system with a 15 frame look-ahead window have shown that the scheme can obtain a 0.6 to 1.6 dB improvement in PSNR over the conventional single encoder on MPEG test video sources. Moreover, measured by the Tektronix Inc.'s Picture Quality Analysis (PQA) System, this scheme achieved 0.87 /spl sim/ 1.35 reduction in PQA measurement. The visual quality improvement is also very evident.
Ligang Lu, Vadim Sheinin
ICASSP (2)2
2003 Real-time MPEG video coding with information look-ahead
abstract
There have been increasing market demands to develop more effective schemes for applying MPEG video coding standard to meet the challenges from the emerging applications. We present a real-time MPEG video coding system with information look-ahead for constant bit rate (CBR) applications, such as video-on-demand (VoD) over ADSL. This scheme employs two MPEG encoders. Both encoders operate at the same CBR. The second encoder has a buffer to delay the input by an amount of time relative to the first encoder to create a look-ahead window. In encoding, the first encoder collects the information of statistics and rate-quality characteristics. An on-line information processor then uses the collected information to derive the best coding strategy for the second encoder to encode the incoming frames in the look-ahead window. The second encoder uses the encoding parameters from the processor as the coding guide to execute the coding strategy and generate the final bitstream. Results on an IBM MPEG-2 encoder testing system with a 15 frame look-ahead window have shown that the scheme can obtain a 0.6 to 1.6 dB improvement in PSNR over the conventional single encoder on MPEG test video sources. Moreover, measured by the Tektronix Inc.'s picture quality analysis (PQA) system, this scheme achieved 0.87 /spl sim/ 1.35 reduction in PQA measurement. The visual quality improvement is also very evident.
Ligang Lu, Vadim Sheinin
ICME2