Vol 8, No 2 (2023)

Streaming Data Processing with Apache Kafka and Spark

Authors:Salim Sheik, Deepak Joshi, Birendar Sirpali

Abstract:Stream processing has become increasingly important in the world of data analytics and real-time decision-making. This paper explores the integration of Apache Kafka and Apache Spark, two popular open-source frameworks, to build robust and scalable streaming data processing pipelines. We discuss the architecture, key components, and use cases of this powerful combination, highlighting its capabilities for handling real-time data at scale.

Keywords:Streaming data processing, Apache Kafka, Apache Spark, Real-time analytics, Data integration, Stream processing, Big data, Use cases, Event-driven architecture, IoT data processing, Fraud detection, Log analysis Data pipelines, Structured Streaming, Machine learning, In-memory computing, Data abstractions, Checkpointing, Fault tolerance

Full Issue

View or download the full issue PDF 55-63

Table of Contents