SIGN IN SIGN UP

Allow TPU `pl.kernel` to be vmapped on nontrivial batch axes.

Also added simple batching rules for `async_copy`, which will issue one DMA call with the singular semaphore.

The primary case to support is pipelining via `emit_pipeline`, including the scalar prefetch cases.

PiperOrigin-RevId: 972088148
I
Ivy Zheng committed
b6143ec1f09d453ac1be4eb08304a0630295813c
Parent: 35d7d10
Committed by jax authors <google-ml-automation@google.com> on 8/27/2026, 7:16:39 PM