Papers
arxiv:2606.21854

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Published on Jun 20
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

ESPnet3 is a modular speech and audio framework that streamlines large-scale experiments through configurable dataset composition, dataset sharding, and lightweight workflow overrides.

Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort to support such experiments. We present ESPnet3, a speech and audio research framework built on a modular system architecture with configuration-driven dataset composition and unified Python-based workflows. ESPnet3 introduces a DataOrganizer abstraction for flexible dataset integration and dataset sharding for memory-efficient large-scale training, while allowing recipe-specific logic through lightweight stage overrides. In OWSM pre-training experiments, ESPnet3 reduces per-epoch training time by 21.1 minutes compared to ESPnet2 and achieves >80\% GPU utilization in multi-node training. Fine-tuning experiments show that new models and datasets can be integrated with around 46 lines of additional code. ESPnet3 will be publicly released with model checkpoints and training logs.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.21854
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.21854 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.21854 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.