modulator / README.md
keshavbhandari's picture
Update README.md
c474ad3 verified
|
Raw
History Blame Contribute Delete
6.23 kB
---
datasets:
- loubb/aria-midi
- projectlosangeles/maestro-v3.0.0
language:
- en
tags:
- music_infilling
- modulation
- music_generation
---
# ๐ŸŽน Change is Key: A Generative Framework for Controllable Musical Modulations
[![Paper](https://img.shields.io/badge/Coming_Soon-blue.svg)](#)
[![Demo](https://img.shields.io/badge/Demo-Audio_%26_Video-green.svg)](https://irrazzib.github.io/modulation-jnmr/)
[![HuggingFace](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model%20Weights-yellow.svg)](https://huggingface.co/keshavbhandari/modulator)
This repository contains the official implementation for **"Change is Key: A Generative Framework for Controllable Musical Modulations"**.
The framework formulates musical modulation as a generative Fill-In-The-Middle (FIM) task using a LLaMA-based architecture. It allows users to select a structural turning point in a MIDI composition, transpose subsequent material to a target key, and autonomously or interactively synthesize a smooth, musically coherent transition bridge.
## **โš™๏ธ Installation**
First, clone this repository and navigate into the modulator folder:
git clone https://github.com/keshavbhandari/modulator.git
cd modulator
Set up your environment using Conda:
module load cuda/12.4.0-gcc-12.2.0 \# Optional: For HPC clusters
conda create \-n modulate python=3.11 \-y
conda activate modulate
conda install pytorch torchvision torchaudio pytorch-cuda=12.4 \-c pytorch \-c nvidia \-y
Install the project and its dependencies:
pip install \-e .
## **๐Ÿš€ Quickstart & Usage**
The primary script for users is modulate\_midi.py. It automatically fetches our pre-trained model weights directly from [Hugging Face (keshavbhandari/modulator)](https://huggingface.co/keshavbhandari/modulator).
Run the commands below from the root modulator folder.
### **1\. Autonomous Infilling (No Chord Conditioning)**
The model relies strictly on left/right musical context and temporal boundary tokens to autonomously bridge into the new key.
python modulate/modulate\_midi.py \\
\--midi\_file path/to/input.mid \\
\--output\_path path/to/output.mid \\
\--semitone\_shift 4 \\
\--chord\_mode none
### **2\. Rule-Guided Modulation (Dynamic Chord Baseline)**
Condition generation on a dynamically calculated functional chord progression that back-cycles diatonic chords to resolve smoothly into the target key.
python modulate/modulate\_midi.py \\
\--midi\_file path/to/input.mid \\
\--output\_path path/to/output.mid \\
\--semitone\_shift 5 \\
\--chord\_mode dynamic
### **3\. LLM-Guided Modulation (Groq API)**
Use a Large Language Model (via the Groq API) to plan custom harmonic progression instructions for the generative bridge based on the surrounding harmonic context.
python modulate/modulate\_midi.py \\
\--midi\_file path/to/input.mid \\
\--output\_path path/to/output.mid \\
\--semitone\_shift 2 \\
\--chord\_mode llm \\
\--api\_key YOUR\_GROQ\_API\_KEY
### **4\. Optimal Modulation Boundary Detection**
Automatically analyze the harmony to identify a structural turning point (e.g., resolving to the tonic) rather than cutting at a generic midpoint.
python modulate/modulate\_midi.py \\
\--midi\_file path/to/input.mid \\
\--output\_path path/to/output.mid \\
\--enable\_optimal\_modulation\_point \\
\--gap\_duration 10000
*(Note: \--gap\_duration is specified in milliseconds. 10000 \= 10 seconds).*
### **5\. Manual Timestamp Control (Milliseconds)**
Explicitly define exact start and end boundaries (in milliseconds) for the left context, the generative gap, and the right context.
python modulate/modulate\_midi.py \\
\--midi\_file path/to/input.mid \\
\--output\_path path/to/output.mid \\
\--left\_segment\_start\_time 70000 \\
\--left\_segment\_end\_time 110000 \\
\--right\_segment\_start\_time 117000 \\
\--right\_segment\_end\_time 140000 \\
\--semitone\_shift 7
### **6\. Simple Infilling (Zero Key Change)**
Repair or infill missing segments without transposing the target key by setting the semitone shift to 0\.
python modulate/modulate\_midi.py \\
\--midi\_file path/to/input.mid \\
\--output\_path path/to/output.mid \\
\--semitone\_shift 0 \\
\--chord\_mode none
## **๐ŸŽ›๏ธ Script Arguments Reference**
| Argument | Default | Description |
| :---- | :---- | :---- |
| \--midi\_file | *Required* | Path to the input .mid file. |
| \--output\_path | infilled\_output.mid | Target save location for output MIDI. |
| \--semitone\_shift | 4 | Semitone transposition for target key (0 for simple infilling). |
| \--chord\_mode | none | Conditioning strategy: none, dynamic, or llm. |
| \--api\_key | None | Groq API key (required when \--chord\_mode llm). |
| \--enable\_optimal\_modulation\_point | False | Analyzes harmony to pick the best structural cut point. |
| \--left\_segment\_start\_time | None | Manual left segment start (ms). |
| \--left\_segment\_end\_time | None | Manual left segment end / gap start (ms). |
| \--right\_segment\_start\_time | None | Manual gap end / right segment start (ms). |
| \--right\_segment\_end\_time | None | Manual right segment end (ms). |
| \--left\_segment\_duration | 30000 | Fallback left context duration (ms). |
| \--right\_segment\_duration | 30000 | Fallback right context duration (ms). |
| \--gap\_duration | 5000 | Fallback generated bridge duration (ms). |
| \--temperature | 1.0 | Sampling temperature for the LLaMA model generation. |
## **๐Ÿง  Model Training**
If you would like to continue training the model or train it on your own custom datasets, you can launch the training script using Accelerate:
accelerate launch \--num\_processes 1 /data/home/acw769/modulator/modulate/models/infiller\_llama/train.py
## **๐Ÿ“œ Citation**
If you use this code, model weights, or evaluation framework in your research, please cite our paper:
@article{bhandari2026change,
title={Change is Key: A Generative Framework for Controllable Musical Modulations},
author={Bhandari, Keshav and Bizzarri, Matteo and Wiggins, Geraint A. and Colton, Simon},
year={2026}
}