Scientific Books

Apache Hudi - The Definitive Guide: Building Robust, Open, And High-performing Data Lakehouses Shiyan Xu O'reilly Media

Overcome challenges in building transactional guarantees on rapidly changing data by using Apache Hudi. With this practical guide, data engineers, data architects, and software architects will...

Overcome challenges in building transactional guarantees on rapidly changing data by using Apache Hudi. With this practical guide, data engineers, data architects, and software architects will discover how to seamlessly build an interoperable lakehouse from disparate data sources and deliver faster insights using their query engine of choice.

Authors Shiyan...

See full description See full description
70 48
Delivery Thu, 13 Aug - Wed, 19 Aug
14,00 €   shipping cost
Sent from Greece
From Toybox 4.7 (29)
Greece
10 pieces
See Books on the page of Toybox

Description

Description

Overcome challenges in building transactional guarantees on rapidly changing data by using Apache Hudi. With this practical guide, data engineers, data architects, and software architects will discover how to seamlessly build an interoperable lakehouse from disparate data sources and deliver faster insights using their query engine of choice.

Authors Shiyan Xu, Prashant Wason, Sudha Saktheeswaran, and Rebecca Bilbro provide practical examples and insights to help you unlock the full potential of data lakehouses for different levels of analytics, from batch to interactive to streaming. You'll also learn how to evaluate storage choices and leverage built-in automated table optimizations to build, maintain, and operate production data applications.

This book helps you:

  • Understand the need for transactional data lakehouses and the challenges associated with building them
  • Get up to speed with Apache Hudi and learn how it makes building data lakehouses easy
  • Explore data ecosystem support provided by Apache Hudi for popular data sources and query engines
  • Perform different write and read operations on Apache Hudi tables and effectively use them for various use cases, including batch and stream applications
  • Implement data engineering techniques to operate and manage Apache Hudi tables
  • Apply different storage techniques and considerations, such as indexing and clustering to maximize your lakehouse performance
  • Build end-to-end incremental data pipelines using Apache Hudi for faster ingestion and fresher analytics

Pages: 350, Dimensions: 17.8x17.8cm

Manufacturer

See full description

Specifications

Specifications

Publisher
O'Reilly Media
Type
Technology, Construction & Building Works, Computers - Informatics
Language
English
Subtitle
-
Cover
Soft
Number of Pages
350
Release Date
11/2025
Publication Date
2025
Dimensions
-
ISBN-13
9781098173838

Important information

Specifications are collected from official manufacturer websites. Please verify the specifications before proceeding with your final purchase. If you notice any problem you can report it here.

See all specifications

Description & Specifications

Overcome challenges in building transactional guarantees on rapidly changing data by using Apache Hudi. With this practical guide, data engineers, data architects, and software architects will discover how to seamlessly build an interoperable lakehouse from disparate data sources and deliver faster insights using their query engine of choice.

Authors Shiyan Xu, Prashant Wason, Sudha Saktheeswaran, and Rebecca Bilbro provide practical examples and insights to help you unlock the full potential of data lakehouses for different levels of analytics, from batch to interactive to streaming. You'll also learn how to evaluate storage choices and leverage built-in automated table optimizations to build, maintain, and operate production data applications.

This book helps you:

  • Understand the need for transactional data lakehouses and the challenges associated with building them
  • Get up to speed with Apache Hudi and learn how it makes building data lakehouses easy
  • Explore data ecosystem support provided by Apache Hudi for popular data sources and query engines
  • Perform different write and read operations on Apache Hudi tables and effectively use them for various use cases, including batch and stream applications
  • Implement data engineering techniques to operate and manage Apache Hudi tables
  • Apply different storage techniques and considerations, such as indexing and clustering to maximize your lakehouse performance
  • Build end-to-end incremental data pipelines using Apache Hudi for faster ingestion and fresher analytics

Pages: 350, Dimensions: 17.8x17.8cm

Manufacturer

Publisher
O'Reilly Media
Type
Technology, Construction & Building Works, Computers - Informatics
Language
English
Subtitle
-
Cover
Soft
Number of Pages
350
Release Date
11/2025
Publication Date
2025
Dimensions
-
ISBN-13
9781098173838

Important information

Specifications are collected from official manufacturer websites. Please verify the specifications before proceeding with your final purchase. If you notice any problem you can report it here.

70,48 €
14,00 €   shipping cost