9798191291697 - ai agent evals: test, measure, and ship reliable llm and agent systems: build evaluation pipelines, catch regressions, and score accuracy, cost, and safety with claude, codex, and python di stallard, smith (4 risultati)

- Brossura
Da: PBShop.store UK, Fairford, GLOS, Regno UnitoPBShop.store UK
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 23,17
EUR 5,85 spedizioneSpedito da Regno Unito a U.S.A.Quantità: Più di 20 disponibili
PAP. Condizione: New. New Book. Shipped from UK. Established seller since 2000.

- Brossura
Da: AHA-BUCH GmbH, Einbeck, GermaniaAHA-BUCH GmbH
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 36,06
EUR 30,50 spedizioneSpedito da Germania a U.S.A.Quantità: 2 disponibili
Taschenbuch. Condizione: Neu. Neuware.

- Brossura
- Print on Demand
Da: California Books, Miami, FL, U.S.A.California Books
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 23,92
Spedizione gratuitaSpedito in U.S.A.Quantità: Più di 20 disponibili
Condizione: New. Print on Demand.

- Brossura
- Print on Demand
Da: CitiRetail, Stevenage, Regno UnitoCitiRetail
Contatta il venditoreVenditore con 5 stelleCondizione: Nuovo
EUR 26,98
EUR 43,11 spedizioneSpedito da Regno Unito a U.S.A.Quantità: 1 disponibili
Paperback. Condizione: new. Paperback. Your Agent Passed the Demo. It's Still Failing in Production. You Just Can't See It Yet.Something changed in how AI gets built - and almost nobody is measuring it correctly.Your agent works in the demo. The refund processes, the ticket closes, the room nods. Then it ships. And somewhere out… in production it's quietly telling customers the wrong policy, burning fifty dollars on a task that should cost forty cents, and leaking one customer's data to another - all while every dashboard glows green. No crash. No error. No alert. Just silent, expensive, trust-destroying failure that you won't discover until a customer does.Here's the uncomfortable truth: evaluating an AI agent is nothing like evaluating a chatbot. A chatbot returns a sentence you can eyeball. An agent produces a trajectory - a chain of tool calls, decisions, and actions where any single step can go wrong in ways a spot-check will never catch. The old playbook of "run it a few times and see if it looks good" isn't just weak. It's the reason your agent is failing right now without your knowledge.This book takes you from "I hope it works" to "I can prove it works - and prove it stays working."You'll build a complete evaluation program from the ground up, following one realistic system the entire way: Northwind Support, a customer-service agent with real tools, real policies, and real failure modes you'll learn to catch by construction. This isn't theory. It's a working project you build chapter by chapter - instrumented traces, a coverage-driven test dataset, code-first graders, calibrated LLM judges, and a CI gate that blocks bad changes before they reach a customer.By the end, you will know how to: Trace what your agent actually did - not what it said it didBuild datasets from real failures that get stronger every time reality attacks themGrade accuracy with honest statistics (no more "82% means nothing" numbers)Measure cost and latency straight off the trace - and know when an 88% agent is a procurement mistakeProbe for jailbreaks and prompt injection continuously, and turn red-team runs into audit evidenceShip model migrations on evidence, not vibes and a leaderboardClose the loop so every production failure becomes a permanent, automated defenseThe window to build this discipline is open right now. A small group of engineers is already treating evals as the core skill of the AI era - gating every deploy, catching regressions in minutes instead of incidents, and shipping agents they can actually defend to a VP, an auditor, or a customer. Everyone else is still running their agent on faith and hoping.Written in a practitioner-to-practitioner voice - no filler, no academic detachment, no AI-sounding fluff - every technique is shown in runnable Python against the same system, so you learn operationally instead of abstractly. Complete appendices give you a 2026 tooling field guide, a critical benchmark reference, copy-ready judge prompts and rubric templates, and a full environment setup so every listing runs as written.Your agent is already failing quietly. The only question is whether you catch it first - or your customers do.Stop shipping on hope. Start shipping on measurement. Build the system that makes your agents provably reliable - starting today. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.