Personal Assistant Benchmark
Scores a personal assistant by what it did on the device
None defined yet.
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
CentificAIResearch is the official Hugging Face organization for Centific Applied AI Research (CAIR).
Centific works with frontier AI labs and enterprises to build production-ready AI systems. We bring together 1.8 million vetted domain experts, 1K+ PhDs, and platforms for data collection, annotation, model fine-tuning, safety evaluation, and localization across 230 languages and locales.
Centific Applied AI Research (CAIR) is focused on one question: what kind of data and evaluation does it take to make AI work reliably in the real world?
We are a team of researchers and engineers working across healthcare AI, physical AI, vision AI, audio AI, AI safety, agentic systems, and multilingual AI.
š See all research publications
We maintain the PRISM Evaluation Suite, covering 7 domains, 12 benchmarks, 25K+ eval tasks, and 50+ models evaluated.
Scores a personal assistant by what it did on the device
Document-work benchmark for healthcare persona
Evaluate AI models on journal entry audit tasks
Co-evolutionary adversarial training demo (DA vs CA)
Explore and compare RL task trajectories