Skip to product information
1 of 1

Hands-On LLM Serving and Optimization

Hands-On LLM Serving and Optimization

Hosting LLMs at Scale

Paperback

Regular price £55.02
Regular price Sale price £55.02

Join our rewards scheme and earn reward points on this purchase!

Earn points on this!

Sign in or Sign up!
View full details
  • Release Date: 26/05/2026
  • Barcode: 9798341621497
  • Genre: Computing & The Internet
  • Sub-Genre: Computer Science
  • Imprint: O'Reilly Media
  • Publisher: O'Reilly Media
Hands-On LLM Serving and Optimization

Hands-On LLM Serving and Optimization

Collapsible content

DESCRIPTION

Hosting LLMs at Scale
As the demand for real-time AI applications grows, along comes this comprehensive guide to the complexities of deploying and optimizing LLMs at scale. The authors take a real-world approach backed by practical examples and code, and assemble essential strategies for designing infrastructures that are equal to the demands of modern AI applications.

Large language models (LLMs) are rapidly becoming the backbone of AI-driven applications. Without proper optimization, however, LLMs can be expensive to run, slow to serve, and prone to performance bottlenecks. As the demand for real-time AI applications grows, along comes Hands-On Serving and Optimizing LLM Models, a comprehensive guide to the complexities of deploying and optimizing LLMs at scale.

In this hands-on book, authors Chi Wang and Peiheng Hu take a real-world approach backed by practical examples and code, and assemble essential strategies for designing robust infrastructures that are equal to the demands of modern AI applications. Whether you're building high-performance AI systems or looking to enhance your knowledge of LLM optimization, this indispensable book will serve as a pillar of your success.

  • Learn the key principles for designing a model-serving system tailored to popular business scenarios
  • Understand the common challenges of hosting LLMs at scale while minimizing costs
  • Pick up practical techniques for optimizing LLM serving performance
  • Build a model-serving system that meets specific business requirements
  • Improve LLM serving throughput and reduce latency
  • Host LLMs in a cost-effective manner, balancing performance and resource efficiency


DELIVERY & RETURNS

UK Delivery:

  • Free delivery on all orders of £10 or more.
  • £1.49 delivery fee on orders below £10.
  • UK orders are shipped via Royal Mail 2nd Class.

International Delivery:

  • Flat rate delivery charges vary by country.

Dispatch and Delivery Times:

  • All orders are shipped from our warehouse in Northampton, UK within 48 hours of receipt during working hours.
  • UK mainland orders typically arrive within 3-5 working days via Royal Mail 2nd Class.
  • International estimated delivery times:
  • Europe & Channel Islands: 7 to 10 working days
  • USA: 7 to 15 working days
  • Rest of the World: 9 to 21 working days

View our full delivery infomation here.

  • OVER

    2 MILLION PRODUCTS

  • 60 MILLION CUSTOMERS

    ACROSS 190 COUNTRIES