张量智推
原名:tensorrt-llm
High-throughput LLM inference on NVIDIA GPUs.
- 分类
- 开发提效
- 版本
- v1.0.1
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-15
- TRACE 评分
- 3.2 / 5
内容概览
NVIDIA's open-source library for optimizing LLM inference with high performance on NVIDIA GPUs. Use TensorRT-LLM when: - Deploying on NVIDIA GPUs (A100, H100, GB200) - Need maximum throughput (24,000+ tokens/sec on Llama 3) - Require low latency for real-time applications - Working with quantized models (FP8, INT4, FP4) - Scaling across multiple GPUs or nodes Use vLLM instead when: - Need simpler setup and Python-first API - Want PagedAttention without TensorRT compilation - Working with AMD GPUs or non-NVIDIA hardware Use llama.cpp instead when: - Deploying on CPU or Apple Silicon - Need edge deployment without NVIDIA GPUs - Want simpler GGUF quantization format - In-flight batching : Dynamic batching during generation - Paged KV cache : Efficient memory management - Flash Attention : Optimized attention kernels - Quantization : FP8, INT4, FP4 for 2-4× faster inference - CUDA graphs : R…