Serverless ETL and Analytics with AWS Glue (Paperback)
Lingua: inglese
Editore: Packt Publishing Limited, Birmingham, 2026
- Brossura
- Nuovo

Da: AussieBookSeller, Truganina, VIC, AustraliaAussieBookSeller
Venditore AbeBooks dal 22 giugno 2007
Condizione: Nuovo
EUR 82,40
Quantità: 1 disponibili
Aggiungi al carrelloDescrizione dell’articolo da parte del venditore
Paperback. Use AWS Glue to integrate growing data sources with serverless ETL, building secure, observable pipelines that support reliable analytics while managing performance and cost across a governed AWS data platform as workloads growKey FeaturesUse runnable code, console walkthroughs, and downloadable examples for core AWS Glue workflowsApply DataOps practices with AWS CDK, Docker, and CI/CD in real-world scenariosLearn from six data specialists with AWS, Spark, Apache Iceberg, and data lake expertiseBook DescriptionWhether you build data pipelines, design cloud architectures, or deliver analytics on AWS, bringing data together is only part of the challenge. You must also keep this data clean, trustworthy, and available while controlling costs. AWS Glue offers serverless data integration, but using it effectively requires decisions about storage, metadata, security, orchestration, monitoring, and performance.This book guides you from modern data management and core AWS Glue features through ingestion from files, streams, SaaS applications, and JDBC sources, preparation, storage layout, metadata, security, sharing, and pipeline operations. Console walkthroughs and runnable examples show how to manage schemas and lineage in AWS Glue Data Catalog, apply AWS Lake Formation access controls, monitor workloads, tune Spark jobs, troubleshoot failures, and manage development with AWS CDK, Docker, and CI/CD. You will also examine analytics, machine learning and generative AI integrations, real-world data lake scenarios, and cost optimization. Learn how Apache Iceberg, Apache Hudi, and Delta Lake add transactions, schema evolution, and efficient data management to data lakes.By the end, you will be able to design, build, operate, and continuously improve a serverless data platform that fits your organization's scale, structure, and priorities.What you will learnDesign scalable serverless ETL pipelines with AWS GlueIngest data from files, streams, SaaS, and JDBC sourcesOptimize file formats, partitions, compression, and layoutsManage schemas, partitions, and lineage in AWS Glue Data CatalogSecure data with access control, encryption, and auditingAutomate testing and multi account CI/CD using AWS CDK and DockerMonitor, tune, and troubleshoot AWS Glue and Spark workloadsApply Apache Iceberg, Hudi, and Delta Lake to data lakes with AWS GlueWho this book is forThis book is for data engineers, ETL developers, cloud architects, and analytics professionals who build or operate data platforms on AWS. It suits readers working on serverless data lakes, Spark ETL, governance, data sharing, reliability, or cost control. It is especially useful if you aim to improve pipeline reliability, governance, or cost visibility as workloads grow. Basic familiarity with the AWS Management Console, Amazon S3, and IAM is recommended. Experience with Python, SQL, or Apache Spark will help with the code examples, and an AWS account is useful for following the walkthroughs. Growing data volumes make fragmented integration harder to govern and scale. This item is printed on demand. Shipping may be from our Sydney, NSW warehouse or from our UK or US warehouse, depending on stock availability.…
Codice articolo 9781835464847
- Titolo
- Serverless ETL and Analytics with AWS Glue (Paperback)
- Autore
- Noritaka Sekiyama
- Editore
- Packt Publishing Limited, Birmingham
- Anno di pubblicazione
- 2026
- Condizione
- new
- Rilegatura
- Paperback
- Lingua
- inglese
- ISBN 10
- 183546484X
- ISBN 13
- 9781835464847
- Edizione
- seconda edizione
Use AWS Glue to integrate growing data sources with serverless ETL, building secure, observable pipelines that support reliable analytics while managing performance and cost across a governed AWS data platform as workloads grow
Key Features
- Use runnable code, console walkthroughs, and downloadable examples for core AWS Glue workflows
- Apply DataOps practices with AWS CDK, Docker, and CI/CD in real-world scenarios
- Learn from six data specialists with AWS, Spark, Apache Iceberg, and data lake expertise
Book Description
Whether you build data pipelines, design cloud architectures, or deliver analytics on AWS, bringing data together is only part of the challenge. You must also keep this data clean, trustworthy, and available while controlling costs. AWS Glue offers serverless data integration, but using it effectively requires decisions about storage, metadata, security, orchestration, monitoring, and performance.
This book guides you from modern data management and core AWS Glue features through ingestion from files, streams, SaaS applications, and JDBC sources, preparation, storage layout, metadata, security, sharing, and pipeline operations. Console walkthroughs and runnable examples show how to manage schemas and lineage in AWS Glue Data Catalog, apply AWS Lake Formation access controls, monitor workloads, tune Spark jobs, troubleshoot failures, and manage development with AWS CDK, Docker, and CI/CD. You will also examine analytics, machine learning and generative AI integrations, real-world data lake scenarios, and cost optimization. Learn how Apache Iceberg, Apache Hudi, and Delta Lake add transactions, schema evolution, and efficient data management to data lakes.
By the end, you will be able to design, build, operate, and continuously improve a serverless data platform that fits your organization's scale, structure, and priorities.
What you will learn
- Design scalable serverless ETL pipelines with AWS Glue
- Ingest data from files, streams, SaaS, and JDBC sources
- Optimize file formats, partitions, compression, and layouts
- Manage schemas, partitions, and lineage in AWS Glue Data Catalog
- Secure data with access control, encryption, and auditing
- Automate testing and multi account CI/CD using AWS CDK and Docker
- Monitor, tune, and troubleshoot AWS Glue and Spark workloads
- Apply Apache Iceberg, Hudi, and Delta Lake to data lakes with AWS Glue
Who this book is for
This book is for data engineers, ETL developers, cloud architects, and analytics professionals who build or operate data platforms on AWS. It suits readers working on serverless data lakes, Spark ETL, governance, data sharing, reliability, or cost control. It is especially useful if you aim to improve pipeline reliability, governance, or cost visibility as workloads grow. Basic familiarity with the AWS Management Console, Amazon S3, and IAM is recommended. Experience with Python, SQL, or Apache Spark will help with the code examples, and an AWS account is useful for following the walkthroughs.
Table of Contents
- Data Management – Introduction and Concepts
- Introduction to Important AWS Glue Features
- Data Ingestion
- Data Preparation
- Data Layouts
- Data Management
- Implementing Resilient Metadata Management
- Data Security
- Data Sharing
- Data Pipeline Management
- Monitoring
- Tuning, Debugging, and Troubleshooting
- Data Analysis
- Machine Learning and Generative AI Integration
- Architecting Data Lakes for Real-World Scenarios and Edge Cases
- Managing End-to-End Development Lifecycle
- Open Table Formats
- Cost Optimization
"Riassunto" può appartenere a un’altra edizione di questo titolo.
Informazioni sull’autore
Noritaka Sekiyama is an experienced big data engineer working at a data and AI company. He is responsible for building scalable data platforms with unified governance in the cloud. He is passionate about software engineering, cloud computing, big data technologies, distributed systems, data platforms, system monitoring, and automation.
Albert Quiroga is a Senior Solutions Architect at Amazon, where he creates solutions and architectural designs for one of the largest data lakes in the world. Prior to that, he spent four years working at AWS, where he specialized in big data technologies such as Amazon EMR, Amazon Athena, AWS Glue, and Amazon SageMaker. His 11 years of experience in the industry have empowered him to work with several Fortune 500 companies to overcome large-scale data and analytics challenges, and he has helped launch and develop features for several AWS services.
Tomohiro Tanaka is a big data specialist with deep, hands-on expertise in data infrastructure. His expertise covers large-scale migrations, performance tuning, and production troubleshooting, with a focus on Apache Spark and Apache Iceberg. He contributes to the Apache Iceberg open-source project and speaks at community events and conferences to help teams adopt Apache Iceberg in practice.
Subramanya Vajiraya is a Senior Cloud Engineer at AWS Sydney specializing in AWS Glue. He obtained his Bachelor of Engineering degree in Information Science & Engineering from NMAM Institute of Technology, Nitte, KA, India, in 2015 and his Master of Information Technology degree in Internetworking from the University of New South Wales, Sydney, Australia, in 2017. He is passionate about helping customers solve challenging technical issues related to their ETL workloads and implement scalable data integration and analytics pipelines on AWS.
Akira Ajisaka is a software engineer with more than 10 years of engineering experience in big data. He enjoys troubleshooting and contributing to OSS.
Ishan Gaur has more than 17 years of IT experience in software development, data engineering, and cloud architecture, building distributed systems and highly scalable data processing pipelines using Apache Spark, Scala, and various AWS data services, such as AWS Glue, Amazon SageMaker Unified Studio, and Amazon EMR. He currently works at AWS as a Principal Cloud Engineer, where he is focused on AI/ML operations and proactive cloud optimization. He works with AWS enterprise customers to design resilient data pipelines, automate incident response, troubleshoot large-scale distributed data platforms, and adopt GenAI-powered services and operational tools. He is passionate about turning reactive support patterns into proactive, self-healing architectures.
"Descrizione articolo" può appartenere a un’altra edizione di questo titolo.
AussieBookSeller
Truganina, VIC, Australia
Venditore AbeBooks dal 22 giugno 2007
Tariffe di spedizione da Australia a U.S.A.
| Articolo | Da 25 a 45 giorni lavorativi | Da 8 a 14 giorni lavorativi |
|---|---|---|
| Primo articolo | EUR 32,51 | EUR 38,66 |
Metodi di pagamento
Informazioni sull’azienda del venditore
The Nile Group Pty Ltd
42 Apex Drive
Truganina, VIC Australia 3029
Condizioni di vendita
We guarantee the condition of every book as it's described on the Abebooks web sites. If you're dissatisfied with your purchase (Incorrect Book/Not as Described/Damaged) or if the order hasn't arrived, you're eligible for a refund within 30 days of the estimated delivery date. If you've changed your mind about a book that you've ordered, please use the Ask bookseller a question link to contact us and we'll respond within 2 business days.
Diritto di recesso
Se sei un consumatore puoi recedere dal contratto in conformità con quanto segue. Per Consumatore si intende qualsiasi persona fisica che agisce per scopi estranei alla propria attività commerciale, imprenditoriale, artigianale o professionale.
Informazioni sul diritto di recesso
Diritto legale di recesso
Hai il diritto di recedere dal presente contratto entro 14 giorni senza fornire alcuna motivazione.
Il periodo di recesso scade dopo 14 giorni dal giorno in cui tu o una terza parte, diversa dal vettore e da te indicata, acquisisce il possesso fisico dell'ultimo bene o dell'ultimo lotto o pezzo.
Per esercitare il diritto di recesso, compila e invia elettronicamente una dichiarazione esplicita sul nostro sito Web, alla voce “I miei acquisti” nella sezione “Mio account”. Ti comunicheremo senza indugio una conferma di ricezione di tale recesso su un supporto durevole (ad es. via e-mail).
Per rispettare il termine di recesso, è sufficiente inviare la comunicazione relativa all'esercizio del diritto di recesso prima della scadenza del periodo di recesso stesso.
Effetti del recesso
In caso di recesso dal presente contratto, ti rimborseremo tutti i pagamenti ricevuti, compresi i costi di spedizione (ad eccezione dei costi supplementari derivanti dalla tua eventuale scelta di un tipo di spedizione diverso dal tipo meno costoso di consegna standard da noi offerto).
Potremo effettuare una detrazione dal rimborso per la perdita di valore dei beni forniti, qualora tale perdita sia il risultato di una manipolazione non necessaria da parte tua.
Eseguiremo il rimborso senza indebito ritardo e non oltre 14 giorni dal giorno in cui saremo informati della tua decisione di recedere dal presente contratto.
Il rimborso sarà effettuato utilizzando lo stesso mezzo di pagamento da te usato per la transazione iniziale, salvo che tu non abbia espressamente concordato altrimenti; in ogni caso, non dovrai sostenere alcun costo quale conseguenza di tale rimborso.
Possiamo trattenere il rimborso finché non avremo ricevuto i beni oppure finché non avrai fornito la prova di averli rispediti, a seconda di quale condizione si verifichi per prima.
Dovrai rispedire i beni o consegnarli a AussieBookSeller, Truganina, Victoria, Australia, senza indebito ritardo e, in ogni caso, entro 14 giorni dal giorno in cui ci hai comunicato la tua volontà di recedere dal presente contratto. Il termine è rispettato se rispedisci i beni prima della scadenza del periodo di 14 giorni. I costi diretti della restituzione dei beni saranno a tuo carico. Sei responsabile solo della diminuzione del valore dei beni risultante da una manipolazione diversa da quella necessaria per stabilire la natura, le caratteristiche e il funzionamento dei beni stessi.
Eccezioni al diritto di recesso
Il diritto di recesso non si applica a:
- La fornitura di giornali, periodici o riviste ad eccezione dei contratti di abbonamento; e
- La fornitura di contenuto digitale non fornito su un supporto materiale (ad es. su un CD o DVD), se al momento dell'invio dell'ordine hai accettato l'inizio dell'esecuzione e hai riconosciuto che non avresti potuto recedere una volta iniziata l'esecuzione.
Condizioni di spedizione
Please note that titles are dispatched from our UK and NZ warehouse. Delivery times specified in shipping terms. Orders ship within 2 business days. Delivery to your door then takes 8-15 days.