Los costos de envío se calcularán en base a esta dirección en todo el sitio.
Selecciona tu país
América
Argentina
Brasil
Canadá
Chile
Colombia
Costa Rica
Ecuador
El Salvador
Estados Unidos
México
Perú
República Dominicana
Uruguay
Europa
Alemania
Austria
Bélgica
Croacia
Dinamarca
Eslovaquia
Eslovenia
España
Finlandia
Francia
Grecia
Hungría
Irlanda
Italia
Letonia
Malta
Noruega
Países Bajos
Polonia
Portugal
Reino Unido
República Checa
Serbia
Suecia
Suiza
Resto del mundo


See, Read, Reason. Building Multimodal AI Applications That Understand Images, Text, and Audio Together (en Inglés)
Richard Boozman (Autor) · Independently published · Tapa Blanda
Quedan más de 100 unidades
$ 118.228The next generation of AI will not understand only text.
It will see images.
Read documents.
Hear audio.
Connect signals across different forms of data.
"See, Read, Reason" is a practical, hands on guide to building multimodal AI applications that can process images, text, and audio together using modern AI models and Python based workflows.
This book shows you how to move beyond single input systems and create applications that reason across multiple modalities.
Why multimodal AI mattersReal world information rarely comes in one format.
Businesses, users, and applications work with:
images and screenshotsdocuments and textvoice recordings and audiovideo frames and metadatamixed data from real environmentsMultimodal AI allows systems to understand these inputs together and produce richer, more useful results.
What you will learnfundamentals of multimodal AI systemshow image, text, and audio models work togetherprocessing visual data for AI applicationsextracting meaning from documents and textworking with speech, audio, and transcriptsdesigning pipelines that combine multiple inputsbuilding reasoning workflows across modalitiesevaluating multimodal model outputsoptimizing latency, cost, and performancedeploying multimodal AI applications in productionFrom separate inputs to unified intelligenceThroughout the book, you will learn how to:
connect vision models with language modelscombine OCR, image understanding, and text reasoningprocess audio into structured insightsbuild assistants that understand mixed inputscreate AI workflows for real world business problemsdesign applications that reason from complete contextEach chapter focuses on practical implementation and product ready patterns.
Practical applicationsdocument intelligence platformsvisual question answering systemsaudio analysis and summarizationcustomer support assistants with image and text inputmeeting intelligence toolsmultimodal research assistantsAI systems for education, healthcare, and business operationsThese examples reflect where modern AI products are heading.
Who this book is forAI engineerssoftware developersdata scientistsproduct buildersstartup foundersprofessionals building next generation AI applicationsIf you want to build AI systems that understand the world more like humans do, this book gives you the roadmap.
See the signal.
Read the context.
Reason across everything.
¿Tienes una pregunta sobre el libro? Inicia sesión para poder agregar tu propia pregunta.

