Premium Only Content
Sparse is Enough in Scaling Transformers (aka Terraformer) | ML Research Paper Explained
#scalingtransformers #terraformer #sparsity
Transformers keep pushing the state of the art in language and other domains, mainly due to their ability to scale to ever more parameters. However, this scaling has made it prohibitively expensive to run a lot of inference requests against a Transformer, both in terms of compute and memory requirements. Scaling Transformers are a new kind of architecture that leverage sparsity in the Transformer blocks to massively speed up inference, and by including additional ideas from other architectures, they create the Terraformer, which is both fast, accurate, and consumes very little memory.
OUTLINE:
0:00 - Intro & Overview
4:10 - Recap: Transformer stack
6:55 - Sparse Feedforward layer
19:20 - Sparse QKV Layer
43:55 - Terraformer architecture
55:05 - Experimental Results & Conclusion
Paper: https://arxiv.org/abs/2111.12763
Code: https://github.com/google/trax/blob/m...
Abstract:
Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propose Scaling Transformers, a family of next generation Transformer models that use sparse layers to scale efficiently and perform unbatched decoding much faster than the standard Transformer as we scale up the model size. Surprisingly, the sparse layers are enough to obtain the same perplexity as the standard Transformer with the same number of parameters. We also integrate with prior sparsity approaches to attention and enable fast inference on long sequences even with limited memory. This results in performance competitive to the state-of-the-art on long text summarization.
Authors: Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, Jonni Kanerva
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
-
LIVE
AgnoLand
3 hours ago🔴 SATURDAY NIGHT OPS | BATTLEFIELD 6 LIVE — PRECISION · CONTROL · CHAOS
105 watching -
4:03:00
TonYGaMinG
6 hours agoARC RAIDERS - DUOS WITH MRR4GER
11.1K -
1:32:57
Jeff Ahern
5 hours ago $11.37 earnedThe Saturday Show with Jeff Ahern
62K25 -
LIVE
Fennis The Gently Devil
2 hours agoTHE BRRRAP PACK: Halo Classic Tournament (commentator's seat)
27 watching -
9:47
MattMorseTV
1 day ago $89.03 earnedDemocrats CAUGHT in $15,000,000 LIE.
146K154 -
1:47:55
Surviving The Survivor: #BestGuests in True Crime
1 day agoDan Markel Murder: Juror from Katie Magbanua's Trial Speaks Out for 1st Time
21.6K -
18:03
stateofdaniel
2 days agoJen Psaki PANICS on Live TV, BACKPEDALS After Smearing Trump with Epstein—Fears LAWSUIT!
28.2K43 -
18:31
Nikko Ortiz
1 day agoKaren You Need A Shower...
46K27 -
1:09:52
VapinGamers
7 hours ago $9.50 earnedTools of the Trade - EP11 Highs and Lows of Streaming with Gothix - !rumbot !music
39.9K2 -
2:14:09
LFA TV
1 day agoRUMBLE RUNDOWN WEEK 6 with JEREMY HERRELL AND SHAWN FARASH 11.15.25 9AM
192K13