TensorRT‑LLM Optimization: Quantization, Kernel Fusion, and Throughput Engineering

Di
- Trex Team
Editore:
- NobleTrex Press

Lingua: Inglese
Formato
Categoria: Non-fiction

"TensorRT‑LLM Optimization: Quantization, Kernel Fusion, and Throughput Engineering"

Built for experienced ML systems engineers, inference specialists, and GPU performance practitioners, this book is a deep guide to making large language models run faster, cheaper, and more predictably with TensorRT‑LLM. Rather than offering generic acceleration advice, it develops a precise mental model of the TensorRT‑LLM stack so readers can understand where performance is won or lost: in quantization choices, graph compilation, fused kernels, KV-cache policy, and serving scheduler behavior.

The book covers the full optimization path from precision strategy and post-training quantization pipelines to engine build configuration, plugin-enabled fusion, attention specialization, and throughput-oriented serving design. Readers will learn how to choose among FP16, BF16, FP8, INT8, and INT4 in hardware-aware ways; validate deployable quantized artifacts; realize fused execution paths in compiled engines; engineer KV-cache behavior for long-context workloads; and benchmark and profile systems with enough rigor to attribute gains to the right layer.

Structured as an advanced, implementation-minded text, the book emphasizes cross-layer tradeoffs rather than isolated tricks. It assumes solid familiarity with transformer inference, CUDA-era GPU concepts, and production deployment concerns, and rewards readers who want durable optimization judgment instead of version-fragile recipes."

Data di uscita

Ebook: 8 maggio 2026

Tag

Scegli il piano che fa per te

Più di 400.000 titoli
Kids Mode (accesso sicuro per bambini)
Scarica e ascolta offline
Disdici quando vuoi

Basic

Le tue prime storie, al prezzo più basso.

6.49 € /mese

1 account
10 ore/mese
Disdici quando vuoi

Prova gratis per 7 giorni

Unlimited

Ascolto illimitato. Dove vuoi, quando vuoi.

9.99 € /mese

1 account
Ascolto illimitato
Disdici quando vuoi

Prova gratis per 14 giorni

Unlimited Annuale

Paghi subito 89.99€/anno, l'equivalente di 7.49€/mese, per 1 anno di ascolto illimitato.

89.99 € /anno

12 mesi al prezzo di 9

1 account
Ascolto illimitato
Disdici quando vuoi

Prova gratis per 14 giorni

Unlimited Family

Risparmia con più account. Ognuno con le proprie storie.

14.99 € /mese

2 account
Ascolto illimitato
Disdici quando vuoi

Prova gratis per 14 giorni