SIGN IN SIGN UP

cuda : add RoPE kernel for mode == 2 (NeoX) (#2760)

* cuda : add RoPE kernel for mode == 2 (NeoX)

* falcon : do not offload the embeddings layer
G
Georgi Gerganov committed
3f460a2b723c8b936ac29ecfd02f244b3adeba55
Parent: 87e3733
Committed by GitHub <noreply@github.com> on 8/25/2023, 8:55:59 AM