relu-revival-normfree

Public

PyTorch implementation of normalization-free LLMs investigating entropic behavior to find desirable activation functions

attention-we entropy-collapse gelu gpt-2 leaky-relu llm-architecture llm-evaluation llm-inference model-optimization normalization-free-training

Creat：2024-10-26T03:34:51

Update：2024-11-03T04:42:01

https://arxiv.org/abs/2410.09637

Stars

Stars Increase

Related projects

Annotated_deep_learning_paper_implementations

Hot

attention

??? 60+ Implementations/tutorials of deep learning papers with side-by-side notes ?; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), ? reinforcement learning (ppo, dqn), capsnet, distillation, ... ?

64723

10个月前

+52today

Vit Pytorch

artificial-intelligence

Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch

24612

10个月前

+24today

Numpy Ml

attention

Machine learning, in numpy

16210

10个月前

+1today

Leedl Tutorial

bert

《李宏毅深度学习教程》（李宏毅老师推荐?，苹果书?），PDF下载地址：https://github.com/datawhalechina/leedl-tutorial/releases

16083

10个月前

+11today

Nlp Tutorial

attention

Natural Language Processing Tutorial for Deep Learning Researchers

14800

1年前

+3today

RWKV LM

attention-mechanism

RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.

14211

10个月前

+11today