Multimodal agents
Connect vision, voice, text, and other inputs into cohesive agent workflows.
Architecting multi-modal and context-aware agents
Learn how to connect reasoning, memory, retrieval, multimodal inputs, and autonomous execution into AI systems that work reliably beyond a demo.


Why I wrote this book
People often ask what it takes to build AI agents that actually work outside of a demo. The answer is learning how to connect reasoning, memory, retrieval, multimodal inputs, and autonomous execution into a system that is reliable in the real world.
This book takes you from your first LangChain workflow to production-ready agentic systems with vision, voice, persistent memory, RAG, multi-hop reasoning, and multi-agent collaboration. Every chapter draws on practical engineering patterns that matter when moving from prototypes to real applications.
I am deeply grateful to everyone who reviewed chapters, shared ideas, challenged assumptions, and helped make the book stronger. I hope it becomes a practical resource you keep coming back to as you build the next generation of AI applications.
Inside the book
Practical patterns for the capabilities that make modern AI applications useful, grounded, and dependable.
Connect vision, voice, text, and other inputs into cohesive agent workflows.
Design context-aware systems that retain useful information across interactions.
Ground responses with retrieval and solve complex tasks through multi-hop reasoning.
Build reliable workflows that plan, select tools, and complete real-world tasks.
Create richer applications that can see, listen, understand, and respond.
Coordinate specialized agents through practical collaboration patterns.
Start building AI agents that move beyond simple chatbots and stand up to real-world requirements.