Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
onset
float64
0.85
2.7k
key_offset
float64
1.09
2.7k
frame_offset
float64
1.09
2.71k
note
int64
21
108
velocity
int64
1
126
6.684375
6.740625
6.740625
105
92
6.685417
6.735417
6.735417
96
87
6.688542
6.730208
6.730208
100
91
6.691667
6.761458
7.81875
69
77
6.692708
6.720833
6.720833
93
93
6.694792
6.751042
7.817708
81
95
6.702083
6.759375
8.382292
76
85
6.71875
6.758333
8.382292
72
64
7.807292
7.861458
8.382292
93
93
7.809375
7.865625
8.382292
105
92
7.817708
7.871875
8.382292
81
89
7.81875
7.876042
8.382292
69
85
7.840625
7.865625
7.998958
104
60
7.998958
8.054167
8.382292
104
79
8.002083
8.032292
8.382292
80
91
8.0125
8.038542
8.382292
68
77
8.017708
8.051042
8.382292
92
87
8.26875
8.522917
9.459375
95
96
8.273958
8.486458
9.222917
100
87
8.273958
8.769792
9.7625
68
68
8.275
8.477083
8.477083
88
97
8.276042
8.536458
9.7625
92
94
8.279167
8.71875
9.7625
71
82
8.280208
8.716667
9.232292
76
99
8.285417
8.685417
9.23125
64
79
8.341667
8.702083
9.7625
74
32
8.344792
8.392708
8.392708
97
21
9.222917
9.277083
9.7625
100
90
9.226042
9.279167
9.714583
88
96
9.23125
9.279167
9.7625
64
72
9.232292
9.276042
9.7625
76
82
9.44375
9.471875
9.7625
60
79
9.45
9.470833
9.7625
84
78
9.455208
9.482292
9.7625
72
53
9.459375
9.483333
9.7625
95
61
9.460417
9.486458
9.7625
96
59
9.685417
9.923958
10.039583
93
103
9.688542
9.897917
10.042708
81
109
9.704167
9.885417
10.036458
57
78
9.70625
9.866667
10.03125
69
86
9.708333
9.925
10.461458
64
77
9.714583
9.759375
9.7625
88
68
9.720833
9.933333
10.461458
60
60
9.780208
9.78125
9.78125
59
19
9.7875
9.839583
10.461458
90
2
9.832292
9.955208
10.461458
88
54
10.03125
10.083333
10.461458
69
82
10.036458
10.090625
10.461458
57
75
10.039583
10.097917
10.461458
93
94
10.042708
10.080208
10.461458
81
97
10.077083
10.119792
10.461458
58
39
10.21875
10.248958
10.461458
68
84
10.219792
10.244792
10.341667
80
92
10.220833
10.257292
10.461458
56
80
10.223958
10.25625
10.461458
92
84
10.341667
10.364583
10.461458
80
4
10.397917
10.569792
10.721875
76
103
10.401042
10.592708
10.720833
88
99
10.401042
10.629167
11.18125
83
96
10.404167
10.588542
10.742708
59
92
10.405208
10.5625
10.720833
64
90
10.4125
10.58125
10.744792
52
82
10.444792
10.48125
10.48125
85
42
10.452083
10.507292
10.507292
55
43
10.720833
10.748958
10.813542
88
92
10.720833
10.759375
11.18125
64
80
10.721875
10.75625
11.18125
76
94
10.728125
10.776042
11.18125
56
65
10.728125
10.784375
11.18125
80
89
10.742708
10.765625
11.18125
59
59
10.744792
10.772917
11.18125
52
49
10.813542
10.814583
11.18125
88
36
10.8875
10.923958
11.18125
48
89
10.891667
10.920833
10.994792
84
95
10.895833
10.929167
11.18125
60
86
10.896875
10.921875
11.18125
72
92
10.994792
11.023958
11.18125
84
33
11.105208
11.328125
11.453125
81
103
11.10625
11.31875
11.445833
69
105
11.115625
11.294792
11.452083
45
85
11.119792
11.30625
11.433333
57
92
11.129167
11.338542
11.98125
52
80
11.163542
11.215625
11.215625
78
33
11.163542
11.330208
11.467708
48
42
11.433333
11.496875
11.98125
57
80
11.445833
11.503125
11.98125
69
91
11.452083
11.495833
11.98125
45
56
11.453125
11.514583
11.98125
81
91
11.459375
11.544792
11.98125
76
80
11.467708
11.495833
11.98125
48
53
11.478125
11.516667
11.98125
46
51
11.648958
11.672917
11.98125
68
99
11.651042
11.677083
11.769792
80
92
11.652083
11.682292
11.98125
56
85
11.6625
11.689583
11.98125
44
76
11.692708
11.734375
11.98125
54
44
11.769792
11.796875
11.98125
80
16
11.916667
12.1875
13.520833
44
47
11.923958
12.136458
12.35
52
94
11.926042
12.21875
12.353125
64
93
End of preview. Expand in Data Studio

PianoVAM v1.2: A Multimodal Piano Performance Dataset

Version History

  • v1.2 (current). Adds Fingering/ (per-note fingering labels for 106 recordings) and Fingering_GT/ (manual fingering annotations for 11 recordings). All other files are unchanged from v1.1.
  • v1.1. metadata.json is the canonical split file. Video files for the 'Sep 04-05' recordings are the sync-corrected versions, and files previously found to have video-MIDI synchronization issues have been relocated across splits; the current splits reflect these corrections.
  • v1.0. Initial release (ISMIR 2025).

Summary

This repository contains the PianoVAM (Video, Audio, Midi and Metadata) dataset, a multi-modal collection of piano performances designed for research in Music Information Retrieval (MIR).

The dataset features synchronized recordings of various piano pieces, providing rich data across several modalities. Our goal is to provide a comprehensive resource for developing and evaluating models that can understand the complex relationship between the visual, auditory, and symbolic aspects of music performance.

How to Cite

If you use the PianoVAM dataset in your research, please cite it as follows:

@inproceedings{kim2025pianovam,
  title={PianoVAM: A Multimodal Piano Performance Dataset},
  author={Kim, Yonghyun and Park, Junhyung and Bae, Joonhyung and Kim, Kirak and Kwon, Taegyun and Lerch, Alexander and Nam, Juhan},
  booktitle={Proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR)},
  year={2025}
}

Usage Guide

0. How to Download the Entire Dataset

The recommended way to download the entire dataset (including all large video files) is to use the huggingface-cli command line tool.

  1. Install the Hugging Face Hub library: If you don't have it installed, open your terminal and run:

    pip install huggingface_hub
    
  2. Download the dataset: Run the following command in your terminal. This will download all repository files, including LFS data, into a folder named PianoVAM_v1.2.

    huggingface-cli download PianoVAM/PianoVAM_v1 --repo-type dataset --local-dir ./PianoVAM_v1.2
    

    (The repo was previously named PianoVAM_v1.0; Hugging Face automatically redirects the old URL to PianoVAM_v1.)

1. Load and Prepare the Dataset

This initial script loads the dataset, constructs the necessary file URLs, and prepares the audio for direct access.

from datasets import load_dataset, Audio
import requests
import os
import numpy as np
from scipy.io.wavfile import write

# Load dataset from the Hub
dataset = load_dataset("PianoVAM/PianoVAM_v1")

# Construct full URLs for all media files
def create_full_media_urls(example):
    base_url = "https://huggingface.co/datasets/PianoVAM/PianoVAM_v1/resolve/main/"
    example["audio_url"] = base_url + example["audio_path"]
    example["video_url"] = base_url + example["video_path"]
    example["midi_url"] = base_url + example["midi_path"]
    return example

dataset = dataset.map(create_full_media_urls)

# Cast the audio_url column for automatic audio decoding
dataset = dataset.cast_column("audio_url", Audio())

# Prepare a sample example from the training set
example = dataset["train"][0]

2. Access Decoded Data

After preparation, you can directly access metadata and the decoded audio array.

# Access metadata
print(f"Piece: {example['piece']} by {example['composer']}")
print(f"Performer: {example['P1_name']}")

# Access the decoded audio data object
audio_data = example["audio_url"]
print(f"Audio sampling rate: {audio_data['sampling_rate']}")
print(f"Audio array shape: {audio_data['array'].shape}")

Output:

Piece: Piano Concerto by E. Grieg
Performer: Yonghyun
Audio sampling rate: 44100
Audio array shape: (32876256,)

3. Download Source Files

Use the following methods to download the original source files to your local machine.

# --- Download Audio File (.wav) ---
audio_array = example["audio_url"]['array']
sampling_rate = example["audio_url"]['sampling_rate']
local_filename = os.path.basename(example['audio_path'])
write(local_filename, sampling_rate, audio_array.astype(np.float32))
print(f"Audio saved as '{local_filename}'")

# --- Download Video File (.mp4) ---
video_url = example['video_url']
local_filename = os.path.basename(video_url)
response = requests.get(video_url)
response.raise_for_status()
with open(local_filename, 'wb') as f:
    f.write(response.content)
print(f"Video saved as '{local_filename}'")

# --- Download MIDI File (.mid) ---
midi_url = example['midi_url']
local_filename = os.path.basename(midi_url)
response = requests.get(midi_url)
response.raise_for_status()
with open(local_filename, 'wb') as f:
    f.write(response.content)
print(f"MIDI saved as '{local_filename}'")

Dataset Description

The dataset consists of various piano pieces performed by multiple pianists. The data was captured simultaneously from a digital piano and high-resolution cameras to ensure precise synchronization between the audio, video, and MIDI streams. The collection is designed to cover a range of musical complexities and styles.

Note on Video Data

Please be aware that all video performances by the pianist named "jiwoo" have had a blur effect applied to the performer's upper body. This was done at the request of the performer to protect their privacy. The keyboard and hands remain fully visible and unaffected.

Directory Structure

The dataset repository is organized into the following directories:

PianoVAM_v1.2/
β”œβ”€β”€ Audio/
β”œβ”€β”€ Fingering/
β”œβ”€β”€ Fingering_GT/
β”œβ”€β”€ Handskeleton/
β”œβ”€β”€ MIDI/
β”œβ”€β”€ TSV/
β”œβ”€β”€ Video/
β”œβ”€β”€ metadata.json
└── README.md

Folder Contents

  • Audio/: Contains the raw audio recordings of the piano performances.

    • Format: Uncompressed WAV (.wav).
    • Sample Rate: 44100 Hz.
  • Video/: Contains the video recordings of the piano performances.

    • Format: MP4 (.mp4).
    • Resolution: 1920x1080 pixels.
    • Frame Rate: 60 fps.
    • Video Codec: H.264 (AVC).
    • Audio Codec: AAC.
  • Handskeleton/: Contains the 3D hand landmark data for each performance.

    • Format: JSON (.json) files.
    • Details: Each file contains frame-by-frame coordinates for the 21 keypoints of both the left and right hands, as captured by MediaPipe Hands.
  • MIDI/: Contains the ground truth performance data recorded directly from a digital piano.

    • Format: Standard MIDI Files (.mid).
    • Details: These files provide the precise timing (onset, offset), pitch, and velocity for every note played.
  • metadata.json: The canonical v1.1 split file. Maps each recording to its data split (train, valid, test, ext-train, special(blurry), special(4hands)) and provides per-recording metadata. Splits reflect the video-MIDI sync corrections.

  • TSV/: Contains pre-processed label data derived from the MIDI files for easier parsing. Each file is a tab-separated value file with 5 columns.

    • Format: TSV (.tsv).
    • Header: onset, key_offset, frame_offset, note, velocity
    • Column Descriptions:
      • onset: The start time of the note in seconds.
      • key_offset: The time when the finger is physically released from the key, in seconds. This is useful for video-based research such as fingering analysis.
      • frame_offset: The time when the sound completely ends, considering pedal usage. This is analogous to the 'offset' used in traditional audio-only transcription.
      • note: The MIDI note number (pitch).
      • velocity: The MIDI velocity (how hard the key was struck).
  • Fingering/: Per-note fingering labels predicted from the video.

    • Format: TSV (.tsv), one file per recording, with the same base name as in TSV/.
    • Header: onset, key_offset, frame_offset, note, velocity, hand, finger
    • Column Descriptions:
      • The first five columns are identical to the file of the same name in TSV/, row for row.
      • hand: L (left), R (right), or Noinfo.
      • finger: 1 (thumb) to 5 (pinky) within that hand, or Noinfo.
    • How the labels were made: Hand landmarks detected with MediaPipe Hands are matched frame by frame to the keys held down in the MIDI. Each note is assigned the finger that stays on its key for most of the note's duration. This is the automatic stage of the fingering method described in Section 5 of the paper. The manual step in the paper, where an annotator resolves ambiguous notes, is not applied to these files.
    • Noinfo labels: A note is marked Noinfo when no finger qualifies or when several fingers qualify and none clearly dominates. Across all 525,483 notes, 19.9% are Noinfo, and most of these are notes where no finger qualifies. The rate varies widely between recordings, from 2.7% to 81.4% (median 15.5%), so please check it for the recordings you use.
    • Accuracy: We compared these labels with the Fingering_GT/ annotations on 1,800 notes from 11 recordings. 88.2% of those notes receive a label. Of the labeled notes, 95.7% have the correct hand and finger, and 99.2% have the correct hand. Most errors are between adjacent fingers. These numbers are measured on the files in this release, so per-recording values can differ from Table 3 of the paper.
    • Coverage: All solo recordings, 106 of the 107 recordings. The four-hands recording 2024-02-15_22-12-41 (split special(4hands)) has no fingering labels because they are provided for solo performances only.
  • Fingering_GT/: Manual fingering annotations, used to evaluate Fingering/.

    • Format: TSV (.tsv) with the same seven columns as Fingering/. Every note has a hand and a finger.
    • Coverage: The first 300 notes of 2024-02-17_22-33-45 and the first 150 notes of 10 other recordings, 1,800 notes in total. Rows are aligned with the first rows of the matching file in TSV/.
    • Source: Annotated by the authors with the GUI annotation tool described in the paper. The same annotations are available as Python lists in fingergt.py in the PianoVAM-Code repository.

    Example:

    import pandas as pd
    
    fing = pd.read_csv("Fingering/2024-02-14_19-10-09.tsv", sep="\t", dtype={"finger": str})
    labeled = fing[fing["hand"] != "Noinfo"]
    print(f"{len(labeled)} of {len(fing)} notes have a fingering label")
    

Planned Updates

  • Improved fingering labels. We are building a new fingering pipeline and plan to release more accurate labels in a future version. They will be added as a new folder. Fingering/ will stay unchanged so that results reported on it remain reproducible.

License

This dataset is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). You are free to share and adapt the material for non-commercial purposes, provided you give appropriate credit and distribute your contributions under the same license.

Contact

For any questions about the dataset, please open an issue in the Community tab of this repository or contact [Yonghyun Kim/yonghyun.kim@gatech.edu].

Downloads last month
1,018

Spaces using PianoVAM/PianoVAM_v1 2