# AI MusiMuse

# 011_COMPOSITION_DATASET.md

Version 1.0

---

# Purpose

The purpose of this document is to define the training data used by the AI Composer.

The Composer never trains directly on WAV files.

Instead, every analyzed track becomes a structured composition sample.

The collection of all composition samples forms the Composition Dataset.

---

# Pipeline

```
Audio

↓

Analysis

↓

MusicDNA

↓

Timeline

↓

Sections

↓

Transitions

↓

Composition Graph

↓

Composition Sample
```

---

# Dataset

The dataset is simply

```
Collection

of

Composition Samples
```

Each sample represents one track.

---

# Composition Sample

Each sample contains

```
Track Metadata

MusicDNA

Timeline

Sections

Transitions

Energy Curve

Density Curve

Texture Timeline

Motif Graph

Style Statistics
```

---

# MusicDNA

Global characteristics

Examples

tempo

key

spectral profile

dynamic profile

rhythm profile

harmony profile

---

# Timeline

Ordered musical events

Example

```
0:00

Pad

0:45

Bass

1:12

Theme

2:20

Break

3:00

Theme Returns
```

---

# Sections

Example

```
Intro

Development

Break

Outro
```

---

# Transition Objects

Example

```
Build

Drop

Fade

Expansion

Reduction
```

Every transition contains

```
from

to

duration

energy_delta

density_delta

confidence
```

---

# Curves

Each track contains

Energy Curve

Density Curve

Brightness Curve

Motion Curve

Harmony Curve

These curves are sampled uniformly.

---

# Motifs

Detected repeating ideas

Example

```
Theme A

↓

Variation

↓

Theme A
```

---

# Statistics

Every sample stores

```
section count

average section duration

average transition duration

maximum density

maximum energy

motif count

repetition ratio
```

---

# Decision History

This is the most important part.

The dataset stores

```
State

↓

Decision

↓

New State
```

Example

```
Energy

0.42

↓

Decision

Add Bass

↓

Energy

0.58
```

---

# Training Record

One composition generates many records.

Example

```
State

↓

Decision
```

```
State

↓

Decision
```

```
State

↓

Decision
```

Thousands of records can be extracted from one song.

---

# Dataset Growth

The dataset grows automatically.

Every analyzed track produces one new sample.

Existing samples are immutable.

---

# Versioning

Every sample stores

```
schema version

MusicDNA version

analyzer versions

builder versions
```

Rebuilding the dataset never overwrites previous versions.

---

# Validation

Every sample must satisfy

deterministic generation

no missing sections

valid timeline

consistent MusicDNA

valid transitions

non-overlapping events

---

# Training Independence

The dataset contains no audio.

Only musical knowledge.

Therefore

different renderers

do not affect training.

---

# Long-Term Goal

The Composition Dataset should become a complete representation of the author's compositional language.

The AI Composer learns from this dataset,

not from raw waveforms.