Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Optional weight tying for Qwen3 and Llama3.2 pretraining (#949)
* optional weight tying for Qwen3 and Llama3.2 * typo
C
casinca committed
9c4be478f89f95414adac173ec034dd777e80974
Parent: e0dbec3
Committed by GitHub <noreply@github.com>
on 1/14/2026, 3:07:04 PM