PaperSwipe

What Distributed Systems Say: A Study of Seven Spark Application Logs

Published 4 years agoVersion 1arXiv:2108.08395

Authors

Sina Gholamian, Paul A. S. Ward

Categories

cs.DCcs.SE

Abstract

Execution logs are a crucial medium as they record runtime information of software systems. Although extensive logs are helpful to provide valuable details to identify the root cause in postmortem analysis in case of a failure, this may also incur performance overhead and storage cost. Therefore, in this research, we present the result of our experimental study on seven Spark benchmarks to illustrate the impact of different logging verbosity levels on the execution time and storage cost of distributed software systems. We also evaluate the log effectiveness and the information gain values, and study the changes in performance and the generated logs for each benchmark with various types of distributed system failures. Our research draws insightful findings for developers and practitioners on how to set up and utilize their distributed systems to benefit from the execution logs.

What Distributed Systems Say: A Study of Seven Spark Application Logs

4 years ago
v1
2 authors

Categories

cs.DCcs.SE

Abstract

Execution logs are a crucial medium as they record runtime information of software systems. Although extensive logs are helpful to provide valuable details to identify the root cause in postmortem analysis in case of a failure, this may also incur performance overhead and storage cost. Therefore, in this research, we present the result of our experimental study on seven Spark benchmarks to illustrate the impact of different logging verbosity levels on the execution time and storage cost of distributed software systems. We also evaluate the log effectiveness and the information gain values, and study the changes in performance and the generated logs for each benchmark with various types of distributed system failures. Our research draws insightful findings for developers and practitioners on how to set up and utilize their distributed systems to benefit from the execution logs.

Authors

Sina Gholamian, Paul A. S. Ward

arXiv ID: 2108.08395
Published Aug 18, 2021

Click to preview the PDF directly in your browser