SIGN IN SIGN UP

🚨 [`Kernels`] Refactor all linear attn models & native kernels fallback (#47630)

* kernel native

* fix warning

* have to fix other modular to be coherent

* remove useless

* use native lib as well

* modular

* combine them

* doc

* ignore kwargs, allow og

* fix

* gdn like paths (missing conv of olmo hybrid)

* olmo hybrid fused conv style

* fix

* fix

* poc mamba2, kernel must compile but torch seems to match

* kernels match --> hf kernels will be needed

* style

* quick fixes

* mamba2 works

* mamba2 suite of models

* oops

* conv1ds across other models

* let's try this

* mamba base implementation

* fixups as per review comments

* fixup mamba tests and other issues

* make no shape check by default (single padded sample should also work)

* remove todo

* propogate mamba1

* style

* fix mambapy + enable on jamba

* zamba1

* style

* fix early cast (leads to non fp32 norm)

* fix padding free path

* fix offload

* bump kernels

* remove todo

* avoid onnx export and fix mamba2 test

* fix bamba test

* update falcon mamba - aligned with all other devices

* oops

* adress review

---------

Co-authored-by: Cyril Vallez <cyril.vallez@gmail.com>
A
Anton Vlasjuk committed
e91f7ef4d3c2b6b96b3a75d6568df581e22a7b84
Parent: 94e2625
Committed by GitHub <noreply@github.com> on 8/5/2026, 9:33:35 AM