Local performance handbook optimizing di tyson ethan (2 risultati)

Autore
Titolo
Perfeziona con la Ricerca avanzata

Perfeziona la tua ricerca

  • Libri (2)

  • Nuovo (2)

a

Fascia di prezzo personalizzata (EUR)

a

  • Lingua: Inglese

    Editore: Independently Published Mai 2026, 2026

    9798195802172

    Serie: Libro 4 di 5 - Architecting Enterprise Agents Series

    • Brossura

    Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 39,62

    EUR 30,50 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 2 disponibili

    Taschenbuch. Condizione: Neu. Neuware - The Local AI Performance Handbook: Optimizing Ollama for Multi-GPU and Hardware AccelerationLocal AI is powerful, but poor configuration can turn expensive hardware into a slow, unstable bottleneck. If your Ollama setup struggles with VRAM limits, weak token throughput, GPU underuse, long context slowdowns, or unreliable multi-user workloads, this handbook gives you the practical performance playbook you need.The Local AI Performance Handbook is a technical guide to building faster, more private, and more reliable Ollama systems across NVIDIA CUDA, AMD ROCm, Apple Silicon, WSL2, Docker, Kubernetes, and multi-GPU environments. It moves beyond basic local model setup and focuses on the engineering details that determine real-world performance: hardware acceleration, VRAM planning, quantization, request concurrency, private RAG, secure deployment, benchmarking, and production maintenance. The book's scope is reflected in its coverage of hardware-specific runtimes, memory engineering, multi-GPU scheduling, quantization, high-concurrency handling, private RAG, deployment, agentic workflows, and troubleshooting.Inside, readers will learn how to: - Configure Ollama for CUDA, ROCm, Apple Silicon, Vulkan, Docker, and WSL2.- Calculate model memory footprints and avoid out-of-memory failures.- Tune VRAM usage, KV cache behavior, context windows, and quantization choices.- Scale Ollama across multiple GPUs and isolate workloads with resource controls.- Benchmark tokens per second, latency, GPU utilization, and system bottlenecks.- Deploy private AI inference with Docker Compose, Kubernetes, health checks, and secure API access.- Build faster private RAG and local agent workflows without depending on cloud APIs.For developers, AI engineers, homelab builders, and technical teams serious about private AI performance, this book turns Ollama from a simple local model runner into a tuned inference platform.

  • Lingua: Inglese

    Editore: Independently published, 2026

    9798195802172

    Serie: Libro 4 di 5 - Architecting Enterprise Agents Series

    • Brossura
    • Print on Demand

    Da: California Books, Miami, FL, U.S.A.California Books

    Venditore con 4 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 22,42

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: Più di 20 disponibili

    Condizione: New. Print on Demand.