Paper: A UNIVERSITY-LEVEL BENCHMARK FOR EVALUATING MATHEMATICAL SKILLS IN LLMS
Toloka
company
Verified
AI & ML interests
Human-expert data for frontier reasoning, safety and agentic AI
Organization Card
Hey, this is Toloka!
models 4
toloka/prompts_reward_model
Text Classification • 82.1M • Updated • 17
toloka/gpt2-large-supervised-prompt-writing
Text Generation • 0.8B • Updated • 87
toloka/gpt2-large-rl-prompt-writing
Text Generation • 0.8B • Updated • 16 • 3
toloka/t5-large-for-text-aggregation
Summarization • Updated • 12 • 7
datasets 14
toloka/homer-v2
Viewer • Updated • 765 • 170 • 5
toloka/HomER
Viewer • Updated • 63 • 27 • 1
toloka/mu-math
Viewer • Updated • 1.08k • 80 • 25
toloka/u-math
Viewer • Updated • 1.1k • 128 • 27
toloka/vist
Viewer • Updated • 39.3k • 90
toloka/VOX-DUB
Viewer • Updated • 7.58k • 220 • 11
toloka/JEEM
Viewer • Updated • 2.2k • 30 • 14
toloka/beemo
Viewer • Updated • 2.19k • 538 • 20
toloka/CLESC
Viewer • Updated • 500 • 14 • 2
toloka/VoxDIY-RusNews
Updated • 104 • 4