Course Hive
Search

Welcome

Sign in or create your account

Continue with Google
or
1.2 Billion Records Per Hour High Performance Kafka and Spark - End to End Data Engineering Project
Play lesson

Google Cloud End to End Data Engineering Projects - 1.2 Billion Records Per Hour High Performance Kafka and Spark - End to End Data Engineering Project

Master End-to-End Data Engineering: Real-Time Streaming, AI Integrations, and High-Performance Systems! Dive into hands-on projects, expert-guided tutorials, and cutting-edge technologies for a standout career in data engineering.

4.0 (2)
25 learners

What you'll learn

Understand and implement real-time streaming with Google Cloud for data engineering projects
Learn how to perform real-time socket streaming using Apache Spark
Master the use of Apache Airflow alongside Spark, Pyspark, Java, and Scala for data engineering
Develop skills to build and optimize high-performance, real-time analytics databases

This course includes

  • 47.5 hours of video
  • Certificate of completion
  • Access on mobile and TV

Summary

Keywords

Full Transcript

PART 2: https://youtu.be/I2K6uCpxL8E Ever wondered how to process 1 billion records per hour seamlessly? In this video, we break down the architecture and tools to make it happen: ✅ Apache Kafka: The backbone of real-time data streaming. ✅ Apache Spark: Lightning-fast processing for massive data pipelines. ✅ ELK Stack: Gain visibility with Elasticsearch, Logstash, and Kibana. ✅ Grafana & Prometheus: Real-time monitoring and performance insights. ✅ Kafka Schema Registry & Control Center: Streamlined management and schema validation. 🎯 What You'll Learn: ✅ How to design a robust architecture for high-throughput data pipelines. ✅ Insights into Python vs. Java Kafka Producers: Which one performs better? ✅ Real-time logging, monitoring, and debugging strategies. 🔥 Why This Matters: If you're in data engineering or want to level up your skills, this video showcases everything you need to build, monitor, and scale an ultra-high-performance streaming platform. Timestamps: 0:00 Introduction 2:31 High Level Architecture Whiteboard 12:55 Data Storage Estimation with workings! 29:33 Clean Architecture 30:39 System Architecture 36:27 System Architecture Setup and Coding 58:21 Python Producer 😩 1:29:27 Java Producer (yay! 😁) 1:33:17 300,000 records per second! 1:36:21 Apache Spark Consumer 2:03:50 Spark Job Optimisation and Statistics 2:15:26 Cluster Health issues 2:15:38 Part 1 Outro 👀 Don't just watch, build it! 🚧 👍 Like, Comment, & Subscribe for more cutting-edge data engineering content! Resources: Full Source Code: https://buymeacoffee.com/yusuf.ganiyu/1-2-billion-records-per-hour-high-performance-kafka-spark-end-end-data Kafka Documentation: https://kafka.apache.org/documentation/ Apache Spark Documentation: https://spark.apache.org/documentation.html #ApacheKafka, #ApacheSpark, #DataEngineering, #BigData, #RealTimeProcessing, #ELKStack, #Grafana, #Prometheus, #KafkaStreams, #BigDataAnalytics, #DataPipeline, #StreamingData, #KafkaMonitoring, #SparkStreaming, #DataArchitecture, #HighPerformanceComputing

Course Hive

Continue this lesson in the app

Install CourseHive on Android or iOS to keep learning while you move.

Related Courses

FAQs

Course Hive
Download CourseHive
Keep learning anywhere