9798183580419 - deep dive into sglang, volume ii: quantization, distributed serving, kernels, and measurement di yang, yin (5 risultati)

Lingua: Inglese
Editore: Independently published, 2026
- Brossura
Da: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 37,29
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

Lingua: Inglese
Editore: Independently published, 2026
- Brossura
Da: PBShop.store UK, Fairford, GLOS, Regno UnitoPBShop.store UK
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 30,13
EUR 7,88 spedizioneSpedito da Regno Unito a U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

Lingua: Inglese
Editore: Independently Published Jun 2026, 2026
- Brossura
Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 48,00
EUR 39,01 spedizioneSpedito da Germania a U.S.A.Quantità: 2 disponibili
Taschenbuch. Condizione: Neu. Neuware - Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang's inference runtime where the core request path meets deployment pressure.This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, d…istributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.The book is written for engineers, researchers, and advanced students who want to understand how modern LLM serving systems preserve correctness while changing representation, placement, execution, and measurement.- How rows stay aligned across optimized execution paths- How KV state moves safely between runtime components- How kernels become eligible or ineligible for a request step- How distributed workers coordinate placement and visibility- How benchmark claims become comparable- How extensions can be evaluated without breaking hidden runtime contractsThis is not a command reference or quick-start guide. It is a deep technical reading of the machinery behind high-throughput inference serving. Each chapter explains the algorithmic or systems problem first, then ties SGLang-specific claims to source files, tests, benchmark artifacts, and operational invariants.Volume II assumes familiarity with the core serving loop from Volume I: admission, prefill and decode scheduling, KV storage, forward execution, sampling, streaming, and cleanup. Readers with GPU systems and transformer inference background can also use the opening preliminaries bridge as a compact orientation before entering the specialized chapters.

Lingua: Inglese
Editore: Independently published, 2026
- Brossura
- Print on Demand
Da: California Books, Miami, FL, U.S.A.California Books
Contatta il venditoreVenditore con 4 stelleCondizione: Nuovo
EUR 30,14
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
Condizione: New. Print on Demand.

Lingua: Inglese
Editore: Independently Published, 2026
- Brossura
- Print on Demand
Da: CitiRetail, Stevenage, Regno UnitoCitiRetail
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 34,21
EUR 43,15 spedizioneSpedito da Regno Unito a U.S.A.Quantità: 1 disponibili
Paperback. Condizione: new. Paperback. Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang's inference runtime where the core request path meets deployment pressure.This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, di…stributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.The book is written for engineers, researchers, and advanced students who want to understand how modern LLM serving systems preserve correctness while changing representation, placement, execution, and measurement.How rows stay aligned across optimized execution pathsHow KV state moves safely between runtime componentsHow kernels become eligible or ineligible for a request stepHow distributed workers coordinate placement and visibilityHow benchmark claims become comparableHow extensions can be evaluated without breaking hidden runtime contractsThis is not a command reference or quick-start guide. It is a deep technical reading of the machinery behind high-throughput inference serving. Each chapter explains the algorithmic or systems problem first, then ties SGLang-specific claims to source files, tests, benchmark artifacts, and operational invariants.Volume II assumes familiarity with the core serving loop from Volume I: admission, prefill and decode scheduling, KV storage, forward execution, sampling, streaming, and cleanup. Readers with GPU systems and transformer inference background can also use the opening preliminaries bridge as a compact orientation before entering the specialized chapters. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.