🚨 [`Kernels`] Refactor all linear attn models & native kernels fallback (#47630)
* kernel native * fix warning * have to fix other modular to be coherent * remove useless * use native lib as well * modular * combine them * doc * ignore kwargs, allow og * fix * gdn like paths (missing conv of olmo hybrid) * olmo hybrid fused conv style * fix * fix * poc mamba2, kernel must compile but torch seems to match * kernels match --> hf kernels will be needed * style * quick fixes * mamba2 works * mamba2 suite of models * oops * conv1ds across other models * let's try this * mamba base implementation * fixups as per review comments * fixup mamba tests and other issues * make no shape check by default (single padded sample should also work) * remove todo * propogate mamba1 * style * fix mambapy + enable on jamba * zamba1 * style * fix early cast (leads to non fp32 norm) * fix padding free path * fix offload * bump kernels * remove todo * avoid onnx export and fix mamba2 test * fix bamba test * update falcon mamba - aligned with all other devices * oops * adress review --------- Co-authored-by: Cyril Vallez <cyril.vallez@gmail.com>
A
Anton Vlasjuk committed
e91f7ef4d3c2b6b96b3a75d6568df581e22a7b84
Parent: 94e2625
Committed by GitHub <noreply@github.com>
on 8/5/2026, 9:33:35 AM