V2N advances visual piano transcription
The Video to Notes system predicts onset, offset, key hold and velocity from piano video.
Why it matters
The work addresses a limitation in audio-based transcription, where sustain pedal effects can obscure physical key release timing. It also broadens visual transcription beyond onset detection toward fuller note-level performance capture.
The key points
- 1.V2N predicts onset, offset, key hold and velocity from video.
- 2.Per-frame multi-task training improved onset accuracy and enabled new outputs.
- 3.The system reports state-of-the-art results on PianoVAM and R3.
Researchers introduced V2N, a visual piano transcription system that converts piano video into note information using a shared temporal backbone with task-specific heads for onset, offset, key hold and velocity. The system is trained with per-frame supervision rather than only at the window center. Reported ablations show multi-task supervision enables offset and velocity prediction while improving onset accuracy, and longer temporal context adds further gains. V2N sets new state-of-the-art results on PianoVAM and R3.
⚡ Try this today
Read the paper before building video-based music transcription systems that need offsets or velocity estimates.
Sources & original reporting
This brief summarizes and links to reporting from the publishers below.
Enjoyed this brief? Get the next one in your inbox.
More in Research
New papers target VLMs' spatial reasoning gap
Researchers propose runtime memory, RL training, benchmarks and 3D generation methods for spatial AI.
New papers probe spatial intelligence in AI vision models
SpaRRTa, SMA and PinpointQA target spatial reasoning gaps in visual and multimodal systems.
Papers refine pseudo-labeling for semi-supervised vision
New arXiv papers target label scarcity in HD mapping and semantic segmentation.