Frozen archive redirect
Mixture of Depths: Dynamically allocating compute in transformer-based language models
This older issue lives in the static archive. If you are not redirected automatically, open the archived issue.
Frozen archive redirect
This older issue lives in the static archive. If you are not redirected automatically, open the archived issue.