A high-throughput and memory-efficient inference and serving engine for LLMs
Choose two branches to see what's changed and to create a pull request.