SIGN IN SIGN UP

Implement non-mapped async IO for CUDA on Windows. (#7896)

* Implement non-mapped async IO for CUDA on Windows. On a fast Gen5 NVMe drive this change improves model load time by >3x while it should be the same (or slightly faster) on any other drive.

* Free resources except for backend.

* Change assertions to exceptions in llama_file, find correct cuda backend to create CUDA resources and respect the use_mmap flag again for CUDA.

* Apply suggestions from code review

Co-authored-by: slaren <slarengh@gmail.com>

* Fix editorconfig and unused variable

* Fix issues with Windows build

---------

Co-authored-by: slaren <slarengh@gmail.com>
M
Markus Tavenrath committed
6a2f0b3474d479bda4ac2ee7cfd5dcdcf0be1f79
Parent: 21be9ca
Committed by GitHub <noreply@github.com> on 6/17/2024, 2:10:15 PM