Episode 398: Apache Kudu with Adar Leiber Dembo

Topics covered
Popular Clips
Episode Highlights
Hadoop Integration
discusses the integration of Apache Kudu within the Hadoop ecosystem, highlighting its compatibility with various components like HDFS and Spark. Kudu is designed to excel in both scans and point lookups, making it versatile for different workloads. explains how Kudu integrates with SQL engines such as Impala and Apache Hive, enhancing its utility in data analytics.
We do try to be really good citizens and integrate with a lot of different systems.
---
This adaptability allows Kudu to fit seamlessly into existing big data workflows, offering a robust solution for analytics 1 2.
Integration Tools
The conversation shifts to the specific tools and services that facilitate Kudu's integration for data processing and analytics. mentions the use of Apache Spark, which provides a high-level framework for data transformations, and Spark SQL for SQL-based queries. These tools, along with others like Presto and Apache Drill, are integral to Kudu's functionality.
Spark offers a much richer and higher level framework for doing data transformations.
---
This integration with various SQL engines allows Kudu to support a wide range of analytical tasks, making it a valuable asset in the big data ecosystem 1 2.
Related Episodes


Episode 157: Hadoop with Philip Zeyliger
Answers 383 questions

Episode 433: Jay Kreps on ksqlDB
Answers 383 questions

Episode 436: Apache Samza with Yi Pan
Answers 383 questions

Episode 193: Apache Mahout
Answers 383 questions

Episode 206: Ken Collier on Agile Analytics
Answers 383 questions

Episode 469: Dhruba Borthakur on Embedding Real-time Analytics in Applications
Answers 383 questions

Episode 393: Jay Kreps on Enterprise Integration Architecture with a Kafka Event Log
Answers 383 questions
SE Radio 560: Sugu Sougoumarane on Distributed SQL Databases
Answers 383 questions

Episode 131: Adrenaline Junkies with DeMarco and Hruschka
Answers 383 questions
Episode 417: Alex Petrov on Database Storage Engines
Answers 383 questions

Episode 179: Cassandra with Jonathan Ellis
Answers 383 questions

Episode 189: Eric Lubow on Polyglot Persistence
Answers 383 questions

Episode 194: Michael Hunger on Graph Databases
Answers 383 questions

Episode 156: Kanban with David Anderson
Answers 383 questions

Episode 229: Flavio Junqueira on Distributed Coordination with Apache ZooKeeper
Answers 383 questions













