SIGN IN SIGN UP

Add Qwen4Exp model (#48337)

* support Qwen4-Exp.

* Fix placement of concatenated checkpoint tensors

* Align Qwen4-Exp checkpoint conversion

* Simplify Qwen4-Exp model integration

* Add Qwen4-Exp parallel coverage and documentation

* revert some equivalent operation

* update qwen4_exp doc.

* fix

* no need to change qwen3.5

* sets activation directly inside norm

* remove specific head models

* native rotary

* simplify

* remove overwrite

* remove non compile criteria

* move outside the class

* remove constants

* simplify config

* indexer is always present

* indexer needs to train

* fix and simplify the selection of valid indices in indexer

* remove external functions

* simplify indexer further

* replace einsum with matmul

* start vectorizing

* simplify

* always use a mask for simplicity

* create and combine mask much more easily in indexer

* simplify hyper connections

* fix all prefixes

* fix ple input by reversing embeddings instead

* capture hidden states

* follow signature order

* always forward both masks

* fix and simplify past context in ngram embedding

* fix

* make do_natural_key more readable

* combine norm and grouped norm

* fix shapes

* fix qsa check

* fix mask

* style

* add a few decorators

* update minimal test

* adapt to newest config

* remove unnecessary test

* combine related test

* fix index mask broadcast error

* fix gdn previous state check when ple is enable

* avoid update conv state in place

* confirmed hybrid_indexed was not used

* fix rope test

* use vlm tester to pass ci

* rewrite layer type in config

* no need to reslice

* short conv up to latest standards

* dilution not supported in kernels 😭

* remove unused attr

* vectorize qsa indexer!!!

* qsa can use a separate rotary for simply

* add qwen_sparse_attention to support block key store!!!

* update test

* modular based on latest main

* make quality happy

* keep simple for now, generalize all library indexers later

* do not keep cos/sin in the cache for memory

* a bit cleaner

* extend 2d position_ids

* ignore rule

* simplify conversion Ops

* fix loading and tests

* do not use attn in tp_plan cause num_kv_head=2

* remove created file

* oupsi

* fix chunked prefill

* fix positions

* fix reverse loading

* improve conversion

* slightly clearer key ordering

* shard the embedding on dim1

* revert test

* force conversion on cpu for device_map

* make things work on simple hardware

* small fix

* oupsi

* latest upstream modular

---------

Co-authored-by: 弥道 <roy.zb@alibaba-inc.com>
Co-authored-by: JJJYmmm <1650675829@qq.com>
C
Cyril Vallez committed
fc5c5bde8e656dad91cbf34e61940d984b1c7b91
Parent: f0e42ef
Committed by GitHub <noreply@github.com> on 8/26/2026, 12:03:40 PM