mmBERT is trained on 3T tokens from over 1800 languages, showing SoTA scores on benchmarks and exceptional low-resource performance
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
DAR: Deontic Reasoning with Agentic Harnesses
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
models 53
jhu-clsp/mmBERT-small
Fill-Mask • Updated • 39.7k • • 86
jhu-clsp/mmBERT-base
Fill-Mask • Updated • 376k • • 244
jhu-clsp/mmBERT-checkpoints
Updated • 4
jhu-clsp/ettin-decoder-1b
Fill-Mask • Updated • 1.25k • 5
jhu-clsp/ettin-decoder-32m
Text Generation • Updated • 635
jhu-clsp/ettin-encoder-1b
Feature Extraction • Updated • 7.93k • 26
jhu-clsp/ettin-encoder-68m
Fill-Mask • Updated • 5.23k • • 5
jhu-clsp/ettin-dec-from-enc-32m
Text Generation • Updated • 15
jhu-clsp/ettin-encoder-150m
Fill-Mask • Updated • 17.9k • • 14
jhu-clsp/ettin-decoder-400m
Text Generation • Updated • 616 • 4
datasets 41
jhu-clsp/cc-temporal-65M
Viewer • Updated • 61.6M • 22
jhu-clsp/SciTaRC
Viewer • Updated • 371 • 58 • 1
jhu-clsp/ManyIH-Bench
Preview • Updated • 88 • 3
jhu-clsp/robust04-instructions
Viewer • Updated • 136k • 512 • 2
jhu-clsp/core17-instructions
Viewer • Updated • 49.4k • 552 • 2
jhu-clsp/news21-instructions
Viewer • Updated • 71.5k • 464 • 1
jhu-clsp/megawika-2
Updated • 8.92k • 6
jhu-clsp/mmBERT-decay-data
Updated • 6.35k • 6
jhu-clsp/mmBERT-midtraining-data
Updated • 4.69k • 1
jhu-clsp/ettin-pretraining-data
Updated • 341k • 10