cuda : add RoPE kernel for mode == 2 (NeoX) (#2760)
* cuda : add RoPE kernel for mode == 2 (NeoX) * falcon : do not offload the embeddings layer
G
Georgi Gerganov committed
3f460a2b723c8b936ac29ecfd02f244b3adeba55
Parent: 87e3733
Committed by GitHub <noreply@github.com>
on 8/25/2023, 8:55:59 AM