Dataset Viewer (First 5GB)
Auto-converted to Parquet Duplicate
uuid
stringlengths
17
17
messages
listlengths
13
588
tools
listlengths
6
6
source
stringclasses
1 value
domain
stringclasses
1 value
Code_Agent_000001
[ { "role": "system", "content": "You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.\n\n## General\n\n- When searching for text or files, prefer using `rg` or `rg --files` respectively because `rg` is much faster than alternatives like `grep`. (If the `rg` ...
[ { "type": "function", "function": { "name": "shell", "description": "Execute a shell command and return its output. Pass `command` as a plain string (recommended); it runs through the user's shell with `-lc`, so standard shell features are supported: pipes (`|`), redirects (`>`, `<`, `>>`), chai...
codex-like-codex.jsonl
Code_Agent
Code_Agent_000002
[ { "role": "system", "content": "You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.\n\n## General\n\n- When searching for text or files, prefer using `rg` or `rg --files` respectively because `rg` is much faster than alternatives like `grep`. (If the `rg` ...
[ { "type": "function", "function": { "name": "shell", "description": "Execute a shell command and return its output. Pass `command` as a plain string (recommended); it runs through the user's shell with `-lc`, so standard shell features are supported: pipes (`|`), redirects (`>`, `<`, `>>`), chai...
codex-like-codex.jsonl
Code_Agent
Code_Agent_000003
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000004
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000005
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000006
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000007
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000008
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000009
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
Code_Agent_000010
[{"role":"system","content":"You are Codex, based on GPT-5. You are running as a coding agent in the(...TRUNCATED)
[{"type":"function","function":{"name":"shell","description":"Execute a shell command and return its(...TRUNCATED)
codex-like-codex.jsonl
Code_Agent
End of preview. Expand in Data Studio

UltraData-SFT-Agent-2609

📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series

English | 中文

📚 Introduction

UltraData-SFT-Agent-2609 is the L3 refined data for Agent instruction-tuning within UltraData's L0-L4 tiered data management framework. Built for the post-training of MiniCPM5-2B, it complements UltraData-SFT-2605 (core-domain SFT) with executable Agent trajectories. The release contains approximately 500,000 samples spanning tool use, Search Agent, Code Agent, and General Agent, covering the full workflow from instruction understanding and tool invocation through environment interaction, verification, error recovery, and final delivery.

The dataset covers task types including skill retrieval and invocation, web search and multi-hop question answering, software engineering and code execution, Office document processing, general function calling, multi-turn memory, financial services, database operations, file reading and writing, and multimodal content processing, together with integrated workflows ranging from skill invocation and file operations to document and office-process execution. Tasks include both single-step or short-chain instruction execution and long-chain tasks that require sustained planning, repeated tool invocation, handling of execution feedback, and intermediate verification, forming a graduated task distribution from basic tool use to complex Agent collaboration.

📢 What's New

  • [2026.09.07] The UltraData-SFT-Agent-2609 dataset is released! Agent instruction-tuning trajectories for the post-training of MiniCPM5-2B, as L3 refined data in the UltraData L0-L4 framework. About 500K samples across tool use, Search Agent, Code Agent, and General Agent. 🚀🚀🚀
  • [2026.09.07] MiniCPM5-2B is released!, the second model in the MiniCPM5 series after MiniCPM5-1B. It is a dense 2B Transformer that scales up the same training recipe, built for on-device, local deployment, and resource-constrained scenarios. It reaches 2B-class open-source SOTA, remains competitive with 4B-class models, and shows particular advantages in coding, mathematics, long-context understanding, tool use, and agentic tasks. UltraData-SFT-Agent-2609 serves as the core Agent SFT dataset for MiniCPM5-2B. 🚀🚀🚀
  • [2026.02.08] The UltraData platform is now live, introducing the L0-L4 tiered data management framework. 🔍🔍🔍

🎯 Dataset Statistics and Capability Coverage

The release contains 483,661 trajectories across four Agent directions. Some General-Agent and Code-Agent tasks were resampled under multiple harnesses, so the counts are trajectories rather than unique tasks.

Direction Samples Share Focus
General-Agent 311,006 64.3% Skill retrieval and invocation, Office / file workflows, multi-harness integrated tasks
Tool-Use 82,760 17.1% Function calling, virtual tools, databases, financial workflows, multi-turn tool use
Code-Agent 69,895 14.5% Software engineering, code edits, tests, and bug fixing across environments
Search-Agent 20,000 4.1% Chinese / English search, multi-hop QA, retrieval planning, web access
Total 483,661 100%

🌟 Dataset Characteristics

  • Complete trajectories: keeps environment observations, tool calls and returns, errors, retries, and final submission—not only the user instruction and the answer.
  • Cross-harness generalization: the same class of task is sampled under different harnesses, tool sets, and system constraints.
  • Turn-level mask: low-quality, redundant, or malformed turns can be excluded from the SFT loss without dropping the whole trajectory.
  • Diagnostic failures retained: some timeout, tool-error, and recovery traces are kept (after de-identification) as auxiliary supervision, not as success demonstrations.

🏗️ Data Construction Pipeline

The dataset is built with a unified pipeline for task generation and Agent trajectory construction:

The figure presents the data-construction process in six steps: task preparation, task rewriting and expansion, environment configuration, cross-harness sampling, outcome verification, and trajectory organization.

  1. Task preparation: task seeds are collected from public tasks, open-source datasets, code repositories, knowledge resources, or manually designed tasks, and are supplemented with roles, goals, constraints, inputs, and delivery requirements.
  2. Task rewriting and expansion: tasks undergo structured rewriting, difficulty grading, and capability labeling in order to generate task variants in different languages, with different task chains and different delivery formats.
  3. Environment configuration: each task is provided with an initial environment, reference files, a tool set, a permission scope, and executable success criteria.
  4. Cross-harness sampling: tasks are placed in different Agent harnesses and in real or simulated environments, where teacher models complete multi-turn interactions and tool calls, environment feedback, and final results are collected.
  5. Outcome verification: results are verified according to task type using program tests, sandbox execution, file-openability checks, format rules, numerical checks, evidence checks, or model review.
  6. Trajectory organization: task instructions, environment observations, assistant behavior, tool inputs, tool returns, errors, and final submissions are organized into a unified message sequence.

🧹 Data Cleaning and Quality Control

Candidate trajectories produced by the pipeline above must pass the following checks before entering the released version. The checks are implemented as a combination of automatic rules, execution verification, and model review, and they span the sampling, verification, and organization stages. They mainly include:

  • Format validation: checks JSONL parseability, message roles, required fields, tool-call format, and parameter schema.
  • Trajectory integrity check: checks whether task instructions, assistant tool calls, tool returns, environment feedback, and final results are correctly paired.
  • Tool behavior check: filters nonexistent tools, illegal parameters, invalid calls, repeated calls, call loops, and tool behavior unrelated to the task.
  • Execution-result verification: verifies whether the task was completed, through sandbox execution, test cases, file artifacts, structured rules, numerical consistency, or evidence tracing.
  • Abnormal-trajectory handling: filters or applies normalized labels to trajectories involving timeouts, environment errors, tool failures, task interruption, and error recovery.
  • Turn-level mask: masks interaction turns that are low-quality, redundant, duplicated, malformed, or otherwise not intended for supervision, so that invalid signals do not enter the SFT loss.
  • Duplicate and contamination handling: performs sample deduplication, near-duplicate detection, cross-split checks, and contamination screening against known evaluation sets.
  • Privacy and safety processing: removes personal information, accounts, keys, cookies, internal paths, and other sensitive fields, and flags or filters dangerous operations, malicious code, and prompt injection.
  • Multi-stage quality review: combines rule-based review, execution-result review, model review, and human spot checks, giving priority to samples with clear goals, executable processes, and verifiable results.

📦 Data Format

Each JSONL line is one trajectory. Released metadata uses five fields: uuid, messages, tools, source, and domain. Tool calls, observations, reasoning, and the final answer live in messages.

{
  "uuid": "agent_sample_000001",
  "messages": [
    {"role": "system", "content": "..."},
    {"role": "user", "content": "..."},
    {"role": "assistant", "content": "...","tool_calls":[...]},
    {"role": "tool", "content": "..."},
    ...
    {"role": "assistant", "content": "..."}
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "",
        "description": "",
        "parameters": {
          "type": "object",
          "title": "parameters",
          "properties": {
            "<parameter_name>": {
              "type": "string",
              "title": "",
              "description": ""
            }
          },
          "required": [""]
        }
      }
    }
  ],
  "source": "UltraData-SFT-Agent-2609",
  "domain": "General_Agent-Office_Agent"
}
Field Type Description
uuid string Sample id.
messages array Chronological system / user / assistant / tool turns.
tools array Tool definitions provided at sampling time; render them into the target chat template when training.
source string Data source.
domain string One of General_Agent-Claw_Agent, General_Agent-Office_Agent, Search_Agent-en, Search_Agent-zh, Search_Agent-plan, Tool_Use, Code_Agent.

Some subsets may keep extra source-specific fields; map them using the subset notes if present.

🚀 Quick Start

from datasets import load_dataset

ds = load_dataset("openbmb/UltraData-SFT-Agent-2609", "Code-Agent", split="train")
ds = load_dataset("openbmb/UltraData-SFT-Agent-2609", "Search-Agent", split="train")
ds = load_dataset("openbmb/UltraData-SFT-Agent-2609", "General-Agent", split="train")
ds = load_dataset("openbmb/UltraData-SFT-Agent-2609", "Tool-Use", split="train")

Available configs: Code-Agent, Search-Agent, General-Agent, Tool-Use.

💡 Intended Uses

  • Agent SFT for MiniCPM-style on-device models, with optional filtering by domain, tool-call density, and trajectory length.
  • Targeted slices: function calling (Tool-Use), search / multi-hop retrieval (Search-Agent), software-engineering agents (Code-Agent), office and skill workflows (General-Agent).
  • Training error recovery and outcome verification, including the retained failure/recovery traces.
  • Mix-ratio studies of Agent data versus core-domain SFT such as UltraData-SFT-2605.

⚠️ Notes and Limitations

  • Static trajectories only: the release does not include environments, tool implementations, tests, or sampling code. Trajectories cannot be replayed or re-accepted locally.
  • Render the sample's tools: tools is frozen from the sampling harness and must be written into the training context; many virtual APIs do not exist in real deployments.
  • Verification ≠ process quality: passing the environment's success check does not mean every step was correct or efficient. Retained failure traces are not success labels.
  • Snapshot freshness: web pages, retrieval hits, files, dependencies, and APIs are frozen at sampling time. Do not treat trajectory facts or versions as up to date.
  • Decontamination scope: screening covered evaluation sets known at construction time; run a new check before introducing a new benchmark.

📜 License and Data Sources

This project is released under the Apache 2.0 license. The dataset aggregates upstream datasets, public repositories, web content, and model-generated trajectories; users must also comply with each upstream license. Apache 2.0 does not override those terms.

Some samples originate from publicly accessible course pages. Public access is not a grant of redistribution, commercial use, or training rights. Rights holders may request removal via the contact channel on the dataset page.

The dataset should not contain plaintext secrets or unauthorized personal data. Please report sensitive content or takedown requests through the same channel.

No unauthorized unchanged redistribution: Without prior written permission from the original authors (or this organization), any institution, organization, or third-party platform is strictly prohibited from directly reposting, mirroring, re-hosting, or commercially repackaging and republishing any artifacts of this project in any form.

📖 Citation

If you find UltraData-SFT-Agent-2609 useful in your research, please consider citing:

@misc{ultradata_sft_agent_2609,
  title        = {UltraData-SFT-Agent-2609},
  author       = {MiniCPM Team},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609}}
}
Downloads last month
7

Models trained or fine-tuned on openbmb/UltraData-SFT-Agent-2609

Space using openbmb/UltraData-SFT-Agent-2609 1

Collections including openbmb/UltraData-SFT-Agent-2609

Paper for openbmb/UltraData-SFT-Agent-2609