Skip to product information
1 of 1

Vision Language Models

Vision Language Models

Building Vlms with Hugging Face

Paperback

Regular price £55.54
Regular price Sale price £55.54

Join our rewards scheme and earn reward points on this purchase!

Earn points on this!

Sign in or Sign up!
View full details
  • Release Date: 23/06/2026
  • Barcode: 9798341624047
  • Genre: Computing & The Internet
  • Sub-Genre: Computer Science
  • Imprint: O'Reilly Media
  • Publisher: O'Reilly Media
Vision Language Models

Vision Language Models

Collapsible content

DESCRIPTION

Building Vlms with Hugging Face
Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farre, Andres Marafioti, and Orr Zohar.

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farre, Andres Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle.

Designed for ML engineers, data scientists, and developers, this guide distills cutting-edge VLM research into practical techniques. Readers will learn how to prepare datasets, select the right architectures, fine-tune and deploy models, and apply them to real-world tasks across a range of industries.

  • Explore core model architectures and alignment techniques
  • Train and fine-tune VLMs with Hugging Face, PyTorch, and others
  • Deploy models for applications like image search and captioning
  • Implement advanced inference strategies, from zero-shot to agentic systems
  • Build scalable VLM systems ready for production use


DELIVERY & RETURNS

UK Delivery:

  • Free delivery on all orders of £10 or more.
  • £1.49 delivery fee on orders below £10.
  • UK orders are shipped via Royal Mail 2nd Class.

International Delivery:

  • Flat rate delivery charges vary by country.

Dispatch and Delivery Times:

  • All orders are shipped from our warehouse in Northampton, UK within 48 hours of receipt during working hours.
  • UK mainland orders typically arrive within 3-5 working days via Royal Mail 2nd Class.
  • International estimated delivery times:
  • Europe & Channel Islands: 7 to 10 working days
  • USA: 7 to 15 working days
  • Rest of the World: 9 to 21 working days

View our full delivery infomation here.

  • OVER

    2 MILLION PRODUCTS

  • 60 MILLION CUSTOMERS

    ACROSS 190 COUNTRIES