Skip to content

A Billion-Parameter Music Model: What 23× More Parameters Bought

Music-LLM · Part 2 · scaling note — 10 October 2026 · 5 min read

Part 2 of 2 The Music-LLM series ← Part 1: Building a Small Transformer That Makes Music

I took a small text-to-music transformer, made it 23 times bigger, trained it on two GPUs, and then measured, as carefully as I could, what actually changed. The honest answer is “less than I hoped, and for a clear reason.”

1.03 B
parameters (45 M before)
10.3 h
to train, 2× RTX A6000
5.362
best validation loss
0.035
nats better than the 45 M model, same clips

1. The model, in one picture

The system has three stages. A frozen T5 encoder turns a prompt such as "jazz music" into vectors. A frozen EnCodec codec turns audio into a grid of integer tokens: 4 codebooks, 2,048 possible values each, 50 frames per second. The part I train is the middle: a transformer decoder that predicts the next audio tokens, looking at the music so far (self-attention) and at the prompt (cross-attention in every block). EnCodec then turns the predicted tokens back into a waveform.

Architecture: text prompt through frozen T5 into a 16-block transformer decoder with cross-attention, four 2048-way heads, decoded to audio by frozen EnCodec; EnCodec tokens feed the decoder's embedding
The 1B configuration: 16 blocks, width 2,048, 16 attention heads. T5 and EnCodec are frozen and are not part of the parameter count.

This is the same recipe as my earlier 45-million-parameter model. Only three numbers change.

2. Making it 23× bigger

The size comes from three settings: the width d, the number of blocks L, and the number of heads. Everything else, the data format, the loss and the code, is identical.

45M1B
width d5122,048
blocks L816
attention heads816
parameters44,876,8001,027,018,752

The parameter count is not an estimate. With K = 4 codebooks, V = 2,048 tokens each, a 1,500-frame context T and a 768-wide T5 output Dt, it is exactly

N = K(V+1)·d  +  T·d  +  L·(14d² + 2·D_t·d + 16d)  +  2d  +  d·K·V
  = 16,785,408 + 3,072,000 + 16 × 61,898,752 + 4,096 + 16,777,216
  = 1,027,018,752

Almost all of it (990 M) is inside the 16 blocks, so doubling the width roughly quadruples the cost.

Why a billion parameters does not fit on one training GPU

Training needs far more memory than the weights. With the AdamW optimizer every parameter carries four float32 numbers: the weight, its gradient and two optimizer statistics, 16 bytes in total. For a billion parameters that is about 16 GB before a single activation is stored.

Training state per model (16 bytes per parameter) 45M 0.72 GB 1B weights 4.1 GB gradients 4.1 GB Adam statistics 8.2 GB 16.4 GB total Split over two 48 GB GPUs, the state fits easily; the activations are what you have to manage.

Four things made the run fit and keep both GPUs busy:

3. Six experiments that all died, and the one wrong number

The first launch tuned six learning rates, and all six died within their first four steps. The cause was the same hardware quirk I wrote about in Part 1: direct GPU-to-GPU copies over PCIe silently corrupt data, and one small cross-GPU operation I had not rerouted through the CPU was still using that path. The loss was a perfectly normal 7.9; a single number was wrong. Every result below was measured after the fix.

4. The training run

The 1B model starts from random weights, so its loss begins at 8.03, just above pure guessing among 2,048 tokens (ln 2048 = 7.62). It trained for 12,000 steps in 10.3 hours. A short learning-rate sweep found that the rate barely matters across a tenfold range, and that the best one is about 5× lower than for the 45 M model: bigger models want smaller steps.

Training curves of the 1B model: loss, per-step loss, learning rate and gradient norm
Validation loss fell at every one of the 30 evaluations, to 5.362.

By the end it had started to memorise: the training audio is only about 206 hours, small for a billion parameters, so the model ran out of new things to learn before it ran out of capacity.

5. Listen

One draw per prompt, 30 seconds each, no re-rolls and no cherry-picking. A lower loss does not by itself mean better-sounding music, so I make no quality claim: judge by ear. Expect genre-flavoured texture, not songs.

blues music
classical music
country music
easy listening music
electronic music
experimental music
folk music
hip-hop music
instrumental music
international music
jazz music
old-time / historic music
pop music
rock music
soul-rnb music

Compressed to mp3 for this page; the original 32 kHz wav files are in the model repository.

6. What the extra parameters bought

The logged validation losses of the two models (5.362 for the 1B and 5.410 for the 45M) were measured on different held-out clips, so they cannot be compared directly. To compare fairly I scored both checkpoints on the same 120 clips that neither model trained on. There the 1B reaches a cross-entropy of 5.320 against 5.355 for the 45M: 23× the parameters bought 0.035 nats. A real gain, but a small one.

Caveat. This scores real music, not how generated audio sounds. A lower loss does not by itself mean better-sounding music, so judge the samples above by ear.

7. What it means, and what comes next

Reproduce and links

pip install torch numpy transformers sentencepiece soundfile huggingface_hub
hf download khashayargh/music-llm-1b-fma --local-dir music-llm-1b-fma
cd music-llm-1b-fma && python example.py "jazz music" --seconds 15

References. FMA: Defferrard et al., ISMIR 2017 (arXiv:1612.01840). MusicGen: Copet et al., NeurIPS 2023 (arXiv:2306.05284). EnCodec: Défossez et al., 2022 (arXiv:2210.13438). T5: Raffel et al., JMLR 2020 (arXiv:1910.10683). GPipe: Huang et al., NeurIPS 2019 (arXiv:1811.06965). The model is for research and education; FMA tracks carry Creative Commons licences, many of them non-commercial.