9798195370404 - parallel computing for ai and ml engineers: build scalable deep learning systems with gpu programming, multi-gpu training, and production workloads di holbrook, m.t (4 risultati)

- Brossura
Da: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 45,46
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

- Brossura
Da: PBShop.store UK, Fairford, GLOS, Regno UnitoPBShop.store UK
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 39,36
EUR 7,90 spedizioneSpedito da Regno Unito a U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

- Brossura
- Print on Demand
Da: California Books, Miami, FL, U.S.A.California Books
Contatta il venditoreVenditore con 4 stelleCondizione: Nuovo
EUR 30,24
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
Condizione: New. Print on Demand.

- Brossura
- Print on Demand
Da: CitiRetail, Stevenage, Regno UnitoCitiRetail
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 43,92
EUR 43,24 spedizioneSpedito da Regno Unito a U.S.A.Quantità: 1 disponibili
Paperback. Condizione: new. Paperback. Stop Guessing. Start Building ML Systems That Actually Scale.Most ML engineers learn GPU computing the hard way - through production failures, mysterious hangs, and models that take three times longer to train than they should. This book gives you the understanding and the tools to get it r…ight the first time.What This Book Covers-GPU architecture internals: CUDA cores, warps, shared memory, and memory coalescing-Writing and optimizing custom CUDA kernels in C++-Data parallel, model parallel, and pipeline parallel training with PyTorch DDP and FSDP-Multi-node training with NCCL, MPI, and InfiniBand-Mixed precision training and gradient scaling-ZeRO optimizer stages 1, 2, and 3 with DeepSpeed-Custom DataLoader optimization and NVIDIA DALI-Production model serving with Triton Inference Server-Kubernetes deployment with GPU autoscaling-Complete profiling workflows with Nsight and PyTorch Profiler-Troubleshooting CUDA OOM, NCCL hangs, and NaN losses-Capacity planning and hardware selection for real workloadsWho This Book Is ForThis book is written for ML engineers, AI researchers, and software engineers working on deep learning infrastructure who want to move beyond single-GPU experiments and build systems that perform at scale. You should be comfortable with Python and have basic familiarity with PyTorch or TensorFlow. No prior CUDA experience required.What Makes This Book DifferentEvery chapter includes complete, runnable code. Architecture diagrams show how components connect. Benchmark results come from real hardware measurements. The troubleshooting appendices address the exact errors that stop real training jobs. This is not a survey of techniques. It is a working engineer's guide to building production parallel ML systems. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.