EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos Paper • 2609.39378 • Published 11 days ago • 74
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? Paper • 2610.07557 • Published 5 days ago • 64
When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents Paper • 2609.32520 • Published 15 days ago • 24
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Paper • 2609.40097 • Published 11 days ago • 49
EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation Paper • 2610.07969 • Published 5 days ago • 19
lmstudio-community/Llama-3-Groq-8B-Tool-Use-GGUF Text Generation • 8B • Updated Jul 18, 2024 • 3.77k • 26
lmstudio-community/Llama-3-Groq-70B-Tool-Use-GGUF Text Generation • 71B • Updated Jul 18, 2024 • 1.2k • 18
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training Paper • 2610.07510 • Published 6 days ago • 14
EVISKILL: Grounding Skill Evolution in Replayable Evidence Paper • 2610.05030 • Published 7 days ago • 52
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks Paper • 2610.08048 • Published 5 days ago • 14
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning Paper • 2610.07043 • Published 6 days ago • 42
LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures Paper • 2610.04292 • Published 8 days ago • 38
Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym Paper • 2609.37267 • Published 12 days ago • 40