Isbn: 9798192442142 - running the gpu fleet: operating gpu clusters for ai with container orchestration, scheduling, observability, and failure management (5 risultati)

Perfeziona la tua ricerca

  • Libri (5)

  • Nuovo (5)

a

Fascia di prezzo personalizzata (EUR)

a

  • Lingua: Inglese

    Editore: Independently published, 2026

    9798192442142

    Serie: Libro 2 di 3 - Scaling AI Systems Series

    • Brossura

    Da: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 14,04

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: Più di 20 disponibili

    PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

  • Lingua: Inglese

    Editore: Independently published, 2026

    9798192442142

    Serie: Libro 2 di 3 - Scaling AI Systems Series

    • Brossura

    Da: PBShop.store UK, Fairford, GLOS, Regno UnitoPBShop.store UK

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 12,99

    EUR 3,88 spedizione 
    Spedito da Regno Unito a U.S.A.

    Quantità: Più di 20 disponibili

    PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

  • Lingua: Inglese

    Editore: Independently Published Aug 2026, 2026

    9798192442142

    Serie: Libro 2 di 3 - Scaling AI Systems Series

    • Brossura

    Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 13,16

    EUR 35,00 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 2 disponibili

    Taschenbuch. Condizione: Neu. Neuware - Keep Your GPU Fleet Fast, Healthy, and Fully UtilizedRunning a modern GPU cluster is a high-stakes challenge. With hardware costs soaring and AI workloads demanding absolute efficiency, cluster operators cannot afford idle time, silent data corruption, or unexpected node failures. Running the GPU Fleet is your definitive, hands-on playbook for operating scale-out AI infrastructure with confidence.Written by experienced cluster engineers, this guide skips the high-level theory to deliver deep, operational blueprints for managing high-performance computing clusters in 2026 and beyond. You will learn how to orchestrate, observe, and maintain accelerators across their entire lifecycle.Inside this comprehensive guide, you will master: - Advanced Kubernetes Orchestration: Configure the NVIDIA GPU Operator, device plugins, MIG profiles, and topology managers without sacrificing predictability.- Efficient Workload Scheduling: Implement gang schedulers like Volcano and Kueue to guarantee atomic, fair-share placements for distributed training.- Deep Observability: Build a robust monitoring stack with DCGM, Prometheus, Grafana, and detailed NCCL profiling to catch errors in minutes.- Failure Management: Troubleshoot frustrating NCCL hangs, InfiniBand link failures, and transceiver degradation using battle-tested playbooks.- Hardware Maintenance & Automation: Detect, isolate, quarantine, and replace faulty nodes across a fleet of thousands of GPUs.Whether you are running Hopper, Blackwell, or next-generation architectures, this book equips you with the exact tools, architectures, and math needed to design multi-tenant clusters that scale seamlessly. Stop wasting compute cycles. Take control of your GPU fleet today.…

  • Condizione: Nuovo

    EUR 14,65

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: Più di 20 disponibili

    Condizione: New. Print on Demand.

  • Lingua: Inglese

    Editore: Independently Published, 2026

    9798192442142

    Serie: Libro 2 di 3 - Scaling AI Systems Series

    • Brossura
    • Print on Demand

    Da: CitiRetail, Stevenage, Regno UnitoCitiRetail

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 16,99

    EUR 43,62 spedizione 
    Spedito da Regno Unito a U.S.A.

    Quantità: 1 disponibile

    Paperback. Condizione: new. Paperback. Keep Your GPU Fleet Fast, Healthy, and Fully UtilizedRunning a modern GPU cluster is a high-stakes challenge. With hardware costs soaring and AI workloads demanding absolute efficiency, cluster operators cannot afford idle time, silent data corruption, or unexpected node failures. Running the GPU Fleet is your definitive, hands-on playbook for operating scale-out AI infrastructure with confidence.Written by experienced cluster engineers, this guide skips the high-level theory to deliver deep, operational blueprints for managing high-performance computing clusters in 2026 and beyond. You will learn how to orchestrate, observe, and maintain accelerators across their entire lifecycle.Inside this comprehensive guide, you will master: Advanced Kubernetes Orchestration: Configure the NVIDIA GPU Operator, device plugins, MIG profiles, and topology managers without sacrificing predictability.Efficient Workload Scheduling: Implement gang schedulers like Volcano and Kueue to guarantee atomic, fair-share placements for distributed training.Deep Observability: Build a robust monitoring stack with DCGM, Prometheus, Grafana, and detailed NCCL profiling to catch errors in minutes.Failure Management: Troubleshoot frustrating NCCL hangs, InfiniBand link failures, and transceiver degradation using battle-tested playbooks.Hardware Maintenance & Automation: Detect, isolate, quarantine, and replace faulty nodes across a fleet of thousands of GPUs.Whether you are running Hopper, Blackwell, or next-generation architectures, this book equips you with the exact tools, architectures, and math needed to design multi-tenant clusters that scale seamlessly. Stop wasting compute cycles. Take control of your GPU fleet today! This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.…