LeRobot documentation

FineART-VLA

You are viewing main version, which requires installation from source. If you'd like regular pip install, checkout the latest stable version (v0.6.1).
Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

FineART-VLA

FineART-VLA builds on π0.5 from Physical Intelligence and its openpi implementation, through LeRobot’s Pi05 policy. It extends this foundation with a trainable PaliGemma language head and a runtime that alternates language generation with action generation. A single checkpoint can predict a low-level subtask, optionally update memory or answer visual questions, and condition its flow-matching action expert on that text.

Use Pi05 when you only need task-conditioned actions. Use FineART-VLA when the policy must generate or consume intermediate language during a rollout.

How FineART-VLA differs from Pi05

CapabilityPi05FineART-VLA
Action modelPaliGemma vision-language prefix + Gemma action expertSame base architecture
Language headNot trained for runtime generationRe-enabled and trained with text cross-entropy
Action conditioningEpisode taskActive low-level subtask plus normalized robot state
Training targetsFlow-matching actionsFlow actions, recipe-selected text, and optional FAST action tokens
Dataset requirementStandard images, state, actions, and taskThe same fields plus language annotations for every language capability you train
RolloutDirect task-to-action policyHierarchical task → subtask → action loop, with optional memory and VQA

FineART-VLA can initialize from a Pi05 checkpoint. The policy architecture remains compatible, while FineART-VLA builds its own processors so recipe labels and FAST labels are not silently replaced by the Pi05 processor stack.

Install

Install LeRobot with the FineART-VLA dependencies:

git clone https://github.com/huggingface/lerobot.git
cd lerobot
python -m venv .venv
source .venv/bin/activate
pip install -e ".[fineart_vla]"

The fineart_vla extra includes the PaliGemma/FAST dependencies. For the optional fused training kernels on Linux with a CUDA GPU, install ".[fineart_vla_kernels]" instead. It adds liger-kernel and the Hugging Face kernels package used by the FlashRT backends. Liger is opt-in with --policy.use_liger_kernels=true: it patches transformers’ PaliGemma RoPE/GeGLU for the whole process, so other PaliGemma-based policies loaded in the same process also run with the fused kernels.

Prepare language-annotated data

FineART-VLA does not infer supervised subtasks from a normal LeRobot dataset during training. The dataset must contain the language targets used by the selected recipe in the optional language_persistent and language_events columns.

At minimum, annotate a continuous subtask timeline so each training frame has an active low-level instruction. Add memory, VQA, interjections, and speech annotations only if the recipe trains those capabilities.

FineART-VLA embeds its default recipe in FineARTVLAConfig, like EO-1. It samples 30% high-level task → subtask prediction and 70% low-level subtask-conditioned action training. Each configuration owns its recipe dictionary, and saved checkpoints carry that dictionary into inference.

Use --policy.recipe_path=/path/to/custom.yaml to override the default. The shared recipes in src/lerobot/configs/recipes/ retain main’s semantics; FineART-VLA does not replace them. For goal-only training, use an inline recipe with one low-level user message containing ${task}, or set recipe=None and disable text/FAST losses for the plain PI05 pipeline.

The default factorizes inference as π(a|o, subtask)·π(subtask|o, task). A custom joint-sequence recipe needs a matching runtime processor that renders the task and generated assistant span with the same causal marks as training. The old policy-side joint prompt builder is no longer used by rollout.

Rollout owns planning cadence (--autosteer_interval_s=3), interactive commands, and synchronous/RTC scheduling. FineART-VLA only generates text when queried; ordinary action calls never generate another subtask behind the runtime’s back.

Combined memory/subtask scratchpads are not the default and require a custom recipe plus a memory-aware runtime controller. Keep memory_scratchpad=false for ordinary subtask generation; older scratchpad checkpoints retain their explicit guard against passing combined responses directly to action control.

Use lerobot-annotate to generate these columns. The repository includes a Hugging Face Jobs launcher that you can edit for your source and destination datasets. For a local annotation run, first install pip install -e ".[annotations]":

HF_TOKEN=hf_... uv run python examples/annotations/run_hf_job.py

Before a long training run, inspect several episodes and verify that subtasks are temporally correct and cover the full demonstration. See Annotation Pipeline for generation and validation, and Language Columns and Recipes for the schema and recipe resolver.

If a dataset has no language columns, recipe rendering becomes a no-op and FineART-VLA falls back to the plain Pi05 prompt path. This is useful for compatibility but does not train the language planner.

Train FineART-VLA

This example initializes FineART-VLA from lerobot/fineart_vla_base, the checkpoint mid-trained on FineART to predict its own next subtask, and trains the embedded 70/30 subtask recipe. --policy.path loads the base checkpoint’s own config, so fine-tuning keeps the FAST token layout (fast_skip_tokens) and the recipe the base was trained with; the other --policy.* flags override it.

lerobot-train \
    --dataset.repo_id=${HF_USER}/my_language_annotated_dataset \
    --policy.path=lerobot/fineart_vla_base \
    --policy.tokenizer_max_length=256 \
    --policy.dtype=bfloat16 \
    --policy.device=cuda \
    --policy.freeze_vision_encoder=false \
    --policy.gradient_checkpointing=true \
    --batch_size=8 \
    --steps=30000 \
    --output_dir=outputs/fineart_vla \
    --job_name=fineart_vla \
    --wandb.enable=true

Start with a small run and confirm that W&B examples show the expected prompt, text target, and action endpoints before scaling up.

To start from the plain PI0.5 checkpoint instead, keep FineART-VLA’s default config and point --policy.pretrained_path at it: --policy.type=fineart_vla --policy.pretrained_path=lerobot/pi05_base. The OpenPI-format weights are converted on load, and the text and FAST losses start from PaliGemma’s language head rather than from FineART mid-training.

Main training controls

OptionDefaultPurpose
policy.recipe_pathNoneOptional override of the embedded 70/30 subtask recipe
policy.text_loss_weight1.0Language-head cross-entropy weight; 0 disables text training
policy.flow_loss_weight10.0Continuous action flow-loss weight
policy.enable_fast_action_losstrueAdds discrete FAST action-token supervision
policy.fast_action_loss_weight1.0FAST cross-entropy weight
policy.knowledge_insulationtrueBlocks action-loss gradients through the VLM K/V path
policy.flow_num_repeats5Reuses one VLM prefix for independent denoising targets
policy.lm_head_lr_scale1.0Scales language-head learning rate; 1.0 uses the base rate
policy.fast_skip_tokens1152FAST id offset; skips <seg>+<loc> so VQA and FAST never collide

fast_skip_tokens=1152 places FAST codes below PaliGemma’s <loc> range. openpi’s pi0-FAST convention is 128 (FAST occupies the <loc> ids); use that value only to stay weight-compatible with checkpoints trained that way, and avoid combining it with the VQA recipe, whose <loc> targets would share embedding rows with FAST codes.

The loss weights are starting points, not dataset-independent constants. Track flow loss and text/FAST losses separately, and inspect generated subtasks rather than selecting a checkpoint from total loss alone.

Dataset-specific FAST tokenizer

By default (--policy.auto_fit_fast_tokenizer=true), training fits a FAST tokenizer on the training dataset, starting from --policy.action_tokenizer_name, and caches it in --policy.fast_tokenizer_cache_dir. This also applies when fine-tuning from a checkpoint such as lerobot/fineart_vla_base, so the FAST targets match the new robot’s action distribution. The fitted tokenizer is saved inside the checkpoint’s processor artifacts, so inference reuses it. Set --policy.auto_fit_fast_tokenizer=false to use action_tokenizer_name unchanged.

Training performance

FineART-VLA uses optimized training paths by default:

  • batches repeated flow targets and suffix projections instead of replaying small operations in Python;
  • caches constant action masks and computes RoPE positions once per forward;
  • selects the text/FAST cross-entropy implementation from target shape and sparsity;
  • skips the mathematically dead VLM/vision backward on knowledge-insulated, flow-only batches;
  • uses native non-reentrant SigLIP layer checkpointing when gradient checkpointing is enabled; and
  • retains the Liger RoPE/GeGLU kernels while avoiding the slower LayerNorm patch at SigLIP shapes.

Optional training backends are disabled by default:

OptionWhen to try it
policy.use_flashrt_adarms=trueFused adaptive RMSNorm and gated residuals on supported CUDA GPUs
policy.use_compiled_text_ce=trueCompiled materialized-logit CE buckets
policy.use_compiled_vision=trueCompiled vision only when the vision pass has no gradients
policy.use_flex_attention=trueProfiled CUDA setups with knowledge insulation and flow_num_repeats > 1; otherwise SDPA is used
policy.use_manual_attention=trueExplicitly profiled shapes where materialized attention is faster
policy.manual_attention_scope=actionRestricts manual attention to action queries

Do not enable every backend blindly. Flex and manual attention are mutually exclusive, and attention/AdaRMS alternatives require knowledge insulation. The benchmark-best configuration used compiled text CE and FlashRT AdaRMS, with Flex/manual attention and compiled vision disabled.

Inference performance

FineART-VLA has two inference loops, and both avoid repeatedly encoding the expensive multimodal prefix:

  1. Action denoising encodes the image/language prefix once, reuses its KV cache across flow steps, constructs action-suffix masks on-device, and crops temporary suffix K/V instead of cloning the prefix cache. It creates the timestep schedule on the device once per chunk, preserving the historical timestep arithmetic. This implementation detail does not establish an end-to-end speedup without benchmarking.
  2. Language decoding uses autoregressive KV caching, so each new token only processes the sampled token against cached image/language keys instead of rerunning the full prefix.

The shared runtime runs language and actions at different rates. Increase --autosteer_interval_s when a subtask remains valid longer, or use /subtask to supply an instruction directly without language generation. A longer interval reduces generation cost but also slows replanning.

Run a checkpoint

Use lerobot-rollout with the robot/camera configuration matching the checkpoint, --interactive=true, and --autosteer_interval_s=3. Choose --inference.type=sync or --inference.type=rtc. In the interactive terminal, /autosteer <goal> enables generated subtasks, /subtask <instruction> supplies one manually, and /start starts robot control. See Interactive language control for complete hardware setup, safety behavior, and recording controls. The former --sim, --mode, and --high_level_hz adapter CLI options are not part of this interface.

Troubleshooting

  • Large flow-loss gradients: FineART-VLA runs KI-off joint prefix/action attention on PyTorch SDPA’s math backend with FP32 attention arithmetic, casting its output back to the model dtype (stock PI0.5 keeps its eager attention). This avoids reduced-precision attention, which can destabilize gradients at low flow timesteps with large query/key activations. Existing checkpoints and normalization layers are unchanged; KI-off still trains through VLM keys and values. Monitor the separate flow_loss, text_loss, and pre-clipping gradient norm rather than relying on the finite total loss alone.
  • No text loss or generated subtasks: confirm the selected recipe can bind the annotations on sampled frames and that policy.text_loss_weight > 0.
  • Subtasks look plausible but actions fail: verify subtask boundaries, normalized state/action statistics, and that low-level recipe samples are present.
  • Text collapses to repeated or location tokens: inspect text-target coverage, language-head learning rate, and the balance between flow, FAST, and text losses.
  • Out of memory: reduce batch size first, then enable gradient checkpointing. Do not enable compiled or alternative attention backends without profiling their memory on your camera count.
  • Slow rollout: separate action latency from language latency, then tune --autosteer_interval_s (how often a new subtask is generated) and --policy.num_inference_steps (flow denoising steps).

Citation

If you use FineART-VLA or the FineART dataset, please cite:

@misc{choghari2026fineart,
  title  = {FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation},
  author = {Choghari, Jade and Kooijmans, Pepijn and Agarwal, Mansi and Ciftci, Yusuf Umut and Doriwala, Aseem and Weaver, Catherine and Sivapurapu, Mouli and Yang, Kai and Lee, Jackson and Wolf, Thomas and Mannam, Pragna},
  year   = {2026},
}

FineART-VLA builds on π₀.₅ and knowledge insulation from Physical Intelligence.

License

This model follows the Apache 2.0 License, consistent with the original OpenPI repository that the PI0.5 backbone is adapted from.

Update on GitHub