Isbn: 9798182198233 - deep dive into sglang, volume i: runtime, scheduling, memory, and decoding (6 risultati)

Perfeziona la tua ricerca

  • Libri (6)

  • Nuovo (6)

a

Fascia di prezzo personalizzata (EUR)

a

  • Lingua: Inglese

    Editore: Independently published, 2026

    9798182198233

    Serie: Libro 3 di 6 - Foundation Books

    • Brossura

    Da: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 38,43

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: Più di 20 disponibili

    PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

  • Lingua: Inglese

    Editore: Independently published, 2026

    9798182198233

    Serie: Libro 3 di 6 - Foundation Books

    • Brossura

    Da: PBShop.store UK, Fairford, GLOS, Regno UnitoPBShop.store UK

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 30,53

    EUR 7,99 spedizione 
    Spedito da Regno Unito a U.S.A.

    Quantità: Più di 20 disponibili

    PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

  • Lingua: Inglese

    Editore: Independently Published Jun 2026, 2026

    9798182198233

    Serie: Libro 3 di 6 - Foundation Books

    • Brossura

    Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 53,67

    EUR 43,73 spedizione 
    Spedito da Germania a U.S.A.

    Quantità: 2 disponibili

    Taschenbuch. Condizione: Neu. Neuware - Deep Dive into SGLang, Volume I explains SGLang's runtime request path as a set of algorithms, data structures, and systems tradeoffs. This volume focuses on the core inference-serving loop: tokenization, admission, prefill and decode scheduling, prefix lookup, KV-cache allocation, forward execution, sampling, streaming, and cleanup. Instead of treating SGLang as a catalog of commands, it builds a working model of the runtime state that makes modern LLM serving possible. Volume I covers: - The inference serving problem and request lifecycle- Transformer inference cost models- SGLang's runtime as a distributed state machine- Continuous batching and chunked prefill- KV-cache memory management- RadixAttention and prefix reuse- Hierarchical caching- Attention backends and forward batches- Model architecture registration- Sampling and logits processing- Structured outputs, reasoning parsers, and tool-call parsing- Speculative decoding- Diffusion language models and blockwise decoding>This is Volume I of Deep Dive into SGLang. Volume II continues into quantization, distributed serving, kernels, hardware backends, benchmarking, correctness, multimodal serving, diffusion inference, and technical extension. Independent explanatory guide. Not affiliated with or endorsed by the SGLang project.…

  • Lingua: Inglese

    Editore: Independently published, 2026

    9798182198233

    Serie: Libro 3 di 6 - Foundation Books

    • Brossura
    • Print on Demand

    Da: California Books, Miami, FL, U.S.A.California Books

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 31,26

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: Più di 20 disponibili

    Condizione: New. Print on Demand.

  • Lingua: Inglese

    Editore: Independently Published, 2026

    9798182198233

    Serie: Libro 3 di 6 - Foundation Books

    • Brossura
    • Print on Demand

    Da: Grand Eagle Retail, Bensenville, IL, U.S.A.Grand Eagle Retail

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 35,15

     Spedizione gratuita 
    Spedito in U.S.A.

    Quantità: 1 disponibile

    Paperback. Condizione: new. Paperback. Deep Dive into SGLang, Volume I explains SGLang's runtime request path as a set of algorithms, data structures, and systems tradeoffs. This volume focuses on the core inference-serving loop: tokenization, admission, prefill and decode scheduling, prefix lookup, KV-cache allocation, forward execution, sampling, streaming, and cleanup. Instead of treating SGLang as a catalog of commands, it builds a working model of the runtime state that makes modern LLM serving possible. Volume I covers: The inference serving problem and request lifecycleTransformer inference cost modelsSGLang's runtime as a distributed state machineContinuous batching and chunked prefillKV-cache memory managementRadixAttention and prefix reuseHierarchical cachingAttention backends and forward batchesModel architecture registrationSampling and logits processingStructured outputs, reasoning parsers, and tool-call parsingSpeculative decodingDiffusion language models and blockwise decodingEach chapter explains the systems problem before implementation details, then ties technical claims back to SGLang source files, tests, or benchmark artifacts. The book is written for readers who already know basic transformer and GPU-systems vocabulary and want a deeper, source-grounded understanding of how a production LLM serving runtime is organized. This is Volume I of Deep Dive into SGLang. Volume II continues into quantization, distributed serving, kernels, hardware backends, benchmarking, correctness, multimodal serving, diffusion inference, and technical extension. Independent explanatory guide. Not affiliated with or endorsed by the SGLang project. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability. …

  • Lingua: Inglese

    Editore: Independently Published, 2026

    9798182198233

    Serie: Libro 3 di 6 - Foundation Books

    • Brossura
    • Print on Demand

    Da: CitiRetail, Stevenage, Regno UnitoCitiRetail

    Venditore con 5 stelle
    Contatta il venditore

    Condizione: Nuovo

    EUR 34,66

    EUR 43,71 spedizione 
    Spedito da Regno Unito a U.S.A.

    Quantità: 1 disponibile

    Paperback. Condizione: new. Paperback. Deep Dive into SGLang, Volume I explains SGLang's runtime request path as a set of algorithms, data structures, and systems tradeoffs. This volume focuses on the core inference-serving loop: tokenization, admission, prefill and decode scheduling, prefix lookup, KV-cache allocation, forward execution, sampling, streaming, and cleanup. Instead of treating SGLang as a catalog of commands, it builds a working model of the runtime state that makes modern LLM serving possible. Volume I covers: The inference serving problem and request lifecycleTransformer inference cost modelsSGLang's runtime as a distributed state machineContinuous batching and chunked prefillKV-cache memory managementRadixAttention and prefix reuseHierarchical cachingAttention backends and forward batchesModel architecture registrationSampling and logits processingStructured outputs, reasoning parsers, and tool-call parsingSpeculative decodingDiffusion language models and blockwise decodingEach chapter explains the systems problem before implementation details, then ties technical claims back to SGLang source files, tests, or benchmark artifacts. The book is written for readers who already know basic transformer and GPU-systems vocabulary and want a deeper, source-grounded understanding of how a production LLM serving runtime is organized. This is Volume I of Deep Dive into SGLang. Volume II continues into quantization, distributed serving, kernels, hardware backends, benchmarking, correctness, multimodal serving, diffusion inference, and technical extension. Independent explanatory guide. Not affiliated with or endorsed by the SGLang project. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. …