Ultimate Multimodal Transformer Models
- 언어학습
- 영어
- 형식
- 컬렉션
논픽션
One Architecture. Infinite Intelligence.
Book Description
Transformer architectures have become the unified foundation of modern AI — powering language models, computer vision systems, and multimodal applications that process text, images, and speech together. Ultimate Multimodal Transformer Models provides a comprehensive, hands-on guide to mastering every major Transformer variant, from foundational encoder-decoder architectures to cutting-edge vision-language models and production GenAI systems.
You begin with the core building blocks of Transformer architecture and text data preparation, then progressively advance through encoder-only models, generative LLMs, RAG, Agentic workflows, and efficient fine-tuning using PEFT, LoRA, and QLoRA. The book then transitions into Vision Transformers, covering ViT, DETR, SAM, CLIP, and Flamingo, before bringing everything together in real-world multimodal applications combining text, vision, and speech using PyTorch and Hugging Face throughout.
By the end of the book, you will be proficient to build, fine-tune, and deploy Transformer-based AI systems across text, vision, and multimodal domains with confidence, applying the right architecture and strategy for every real-world use case!
What you will learn
? Build and deploy Transformer models for text, vision, and multimodal AI tasks.
? Fine-tune large language models efficiently using PEFT, LoRA, and QLoRA techniques.
? Develop production-ready GenAI applications using RAG pipelines and Agentic AI workflows.
? Apply LLMs to real-world NLP tasks including summarization, question answering, and classification.
? Implement Vision Transformers, DETR, and SAM for object detection and image segmentation tasks.
? Integrate multimodal AI systems combining text, vision, and speech using CLIP and Flamingo architectures.
Table of Contents
1. The Rise of Transformer Models in Sequence Learning
2. Text Data Preparation for Transformer Models
3. Building Blocks of Transformer Architecture
4. Encoder-only Transformer Configurations
5. Generative Transformers and LLM Architectures
6. Customizing LLMs Using Retrieval-Augmented Generation (RAG)
7. Efficient Fine-Tuning Techniques with PEFT and LoRA
8. Orchestrating LLMs with Tools and Memory
9. Introduction to Vision Transformer Models
10. Vision Transformers for Image Classification
11. Object Detection and Segmentation with Transformer Architectures
12. Vision-Language Models and Multimodal LLMs
13. Real-World Multimodal GenAI Applications
14. Image Generation with Vision Transformers
15. The Future of GenAI with Transformers
Index
© 2026 Orange Education Pvt Ltd (전자책): 9788169646833
출시일
전자책: 2026년 6월 2일
다른 사람들도 즐겼습니다 ...
- 어린이날이 사라진다고? 노수미
- 그 많던 싱아는 누가 다 먹었을까 박완서
- 신기한 맛 도깨비 식당 1 김용세, 김병선
- who? 마이클 잭슨 툰쟁이, 한나나
- 쓸 만한 인간 박정민
- 기분이 태도가 되지 않게 레몬심리
- 용선생 처음 세계사1: 고대 문명~중세: 고대 문명~중세 김선혜, 정지윤, 노남희, 뭉선생, 윤효식, 이우일, 김선빈, 사회평론 역사연구소
- 돈의 속성 김승호
- who? 이순신 스튜디오청비, 이수겸
- 바람의 화원 1 이정명
- who? 단군·주몽 정병훈, 최재훈
- Henry and Mudge: The First Book Cynthia Rylant
- 어린 왕자 앙투안 드 생텍쥐페리
- 명탐정 셜록 홈즈 1 아서 코난 도일
- 신기한 맛 도깨비 식당 3 김용세, 김병선
언제 어디서나 스토리텔
국내 유일 해리포터 시리즈 오디오북
5만권이상의 영어/한국어 오디오북
키즈 모드(어린이 안전 환경)
월정액 무제한 청취
언제든 취소 및 해지 가능
오프라인 액세스를 위한 도서 다운로드
스토리텔 언리미티드
5만권 이상의 영어, 한국어 오디오북을 무제한 들어보세요
13800 원 /월
계정 1개
무제한 청취
사용자 1인
무제한 청취
언제든 해지하실 수 있어요
패밀리
친구 또는 가족과 함께 오디오북을 즐기고 싶은 분들을 위해
매달 21500 원 원 부터
2-3 개 계정
무제한 청취
2-3 계정
무제한 청취
언제든 해지하실 수 있어요
본인 + 1 가족 구성원
2 개 계정21500 원 /월
