👋 Open to Work
Andrew Thompson
AndrewThompson1233
AI & ML interests
Deep Learning / Model Architecture & Training
Recent Activity
liked a model 2 days ago
AndrewThompson1233/maba-instant-v1 new activity 2 days ago
AndrewThompson1233/maba-v1-architecture:Factorized embeddings intervention for GPT-2 published a model 2 days ago
AndrewThompson1233/maba-instant-v1Organizations
Factorized embeddings intervention for GPT-2
11
#1 opened 6 days ago
by
gpjt
Mitigating chat overfitting on small corpora via block recycling and RoPE
13
#1 opened 18 days ago
by
AndrewThompson1233
Gated MLA context memory scaling and vocabulary tax on 2.8B active compute
👍 1
15
#1 opened 17 days ago
by
AndrewThompson1233
State attenuation across depth and long-context needle retention in 3:1 NoPE hybrids
4
#1 opened 15 days ago
by
AndrewThompson1233
Overcoming the ARC-Easy floor and vocabulary budget on an 80M footprint
❤️ 1
8
#1 opened 18 days ago
by
AndrewThompson1233
The 256-token context cliff and sequence scaling in custom 346M chat architectures
10
#1 opened 15 days ago
by
AndrewThompson1233
The 39% vocabulary parameter tax and layer allocation on Colab T4 budgets
3
#1 opened 15 days ago
by
AndrewThompson1233
The 67M untied vocabulary footprint and context scaling on 6GB RTX 3050 budgets
2
#1 opened 15 days ago
by
AndrewThompson1233
Subspace entanglement in attention o_proj edits and downstream expert routing drift
2
#1 opened 15 days ago
by
AndrewThompson1233