9798195860981 - vllm and high-performance inference: memory optimization, parallel execution, token streaming, and scalable model serving: 2 di cypher, camila (5 risultati)

Lingua: Inglese
Editore: Independently Published, 2026
Serie: Libro 2 di 2 - Large Language Model Refinement and Inference Series
- Brossura
Da: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 22,53
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

Lingua: Inglese
Editore: Independently Published, 2026
Serie: Libro 2 di 2 - Large Language Model Refinement and Inference Series
- Brossura
Da: PBShop.store UK, Fairford, GLOS, Regno UnitoPBShop.store UK
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 20,65
EUR 4,86 spedizioneSpedito da Regno Unito a U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

Lingua: Inglese
Editore: Independently published, 2026
Serie: Libro 2 di 2 - Large Language Model Refinement and Inference Series
- Brossura
- Print on Demand
Da: California Books, Miami, FL, U.S.A.California Books
Contatta il venditoreVenditore con 4 stelleCondizione: Nuovo
EUR 20,45
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
Condizione: New. Print on Demand.

Lingua: Inglese
Editore: Amazon Digital Services LLC - Kdp Mai 2026, 2026
Serie: Libro 2 di 2 - Large Language Model Refinement and Inference Series
- Brossura
Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 25,00
EUR 61,66 spedizioneSpedito da Germania a U.S.A.Quantità: 2 disponibili
Taschenbuch. Condizione: Neu. Neuware - Once a language model has been refined, its effectiveness depends on how well it can be delivered in real-world environments. This book examines the systems and techniques that enable efficient inference, with a particular focus on vLLM and the architectural decisions that support high-thr…oughput execution.The text begins by establishing the relationship between model size, hardware constraints, and response latency. It then explores how memory is managed during inference, including strategies that reduce overhead while maintaining output quality. Concepts such as batching, caching, and token-level scheduling are presented in a way that reveals their practical impact on performance.A central theme of the book is parallel execution, where multiple requests are handled simultaneously without degrading responsiveness. The discussion highlights how modern inference frameworks distribute workloads, coordinate computation, and maintain consistency across concurrent processes.Token streaming is examined as a critical component of user-facing systems, showing how incremental output generation improves perceived responsiveness and interaction flow. The material connects these techniques to broader system considerations, including scaling across machines, managing resource allocation, and maintaining stability under load.As the book progresses, it presents a unified view of inference as both a technical and operational challenge. It demonstrates how decisions made at the system level directly influence user experience, cost efficiency, and reliability.By the end, readers will have a clear understanding of how optimized inference transforms a refined model into a responsive and scalable system capable of operating under demanding conditions.

Lingua: Inglese
Editore: Independently Published, 2026
Serie: Libro 2 di 2 - Large Language Model Refinement and Inference Series
- Brossura
- Print on Demand
Da: CitiRetail, Stevenage, Regno UnitoCitiRetail
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 24,66
EUR 43,24 spedizioneSpedito da Regno Unito a U.S.A.Quantità: 1 disponibili
Paperback. Condizione: new. Paperback. Once a language model has been refined, its effectiveness depends on how well it can be delivered in real-world environments. This book examines the systems and techniques that enable efficient inference, with a particular focus on vLLM and the architectural decisions that support high-thro…ughput execution.The text begins by establishing the relationship between model size, hardware constraints, and response latency. It then explores how memory is managed during inference, including strategies that reduce overhead while maintaining output quality. Concepts such as batching, caching, and token-level scheduling are presented in a way that reveals their practical impact on performance.A central theme of the book is parallel execution, where multiple requests are handled simultaneously without degrading responsiveness. The discussion highlights how modern inference frameworks distribute workloads, coordinate computation, and maintain consistency across concurrent processes.Token streaming is examined as a critical component of user-facing systems, showing how incremental output generation improves perceived responsiveness and interaction flow. The material connects these techniques to broader system considerations, including scaling across machines, managing resource allocation, and maintaining stability under load.As the book progresses, it presents a unified view of inference as both a technical and operational challenge. It demonstrates how decisions made at the system level directly influence user experience, cost efficiency, and reliability.By the end, readers will have a clear understanding of how optimized inference transforms a refined model into a responsive and scalable system capable of operating under demanding conditions. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.