张量智推

原名:tensorrt-llm

High-throughput LLM inference on NVIDIA GPUs.

分类
开发提效
版本
v1.0.1
作者
弈韬(@ra1nzzz)
下载
1
收藏
0
发布
2026-08-25
更新
2026-09-15
TRACE 评分
3.2 / 5

内容概览

NVIDIA's open-source library for optimizing LLM inference with high performance on NVIDIA GPUs. Use TensorRT-LLM when: - Deploying on NVIDIA GPUs (A100, H100, GB200) - Need maximum throughput (24,000+ tokens/sec on Llama 3) - Require low latency for real-time applications - Working with quantized models (FP8, INT4, FP4) - Scaling across multiple GPUs or nodes Use vLLM instead when: - Need simpler setup and Python-first API - Want PagedAttention without TensorRT compilation - Working with AMD GPUs or non-NVIDIA hardware Use llama.cpp instead when: - Deploying on CPU or Apple Silicon - Need edge deployment without NVIDIA GPUs - Want simpler GGUF quantization format - In-flight batching : Dynamic batching during generation - Paged KV cache : Efficient memory management - Flash Attention : Optimized attention kernels - Quantization : FP8, INT4, FP4 for 2-4× faster inference - CUDA graphs : R…

查看 SKILL 详情