# AI MusiMuse

# 007_MUSIC_REPRESENTATION.md

Version 1.0

---

# Purpose

This document defines the internal representation of music used by AI MusiMuse.

The system does not learn directly from raw audio.

Instead, every composition is transformed into multiple abstraction layers.

Each layer answers a different musical question.

---

# Representation Layers

Every track is represented as seven independent layers.

```
Track
│
├── Metadata
├── Audio
├── MusicDNA
├── Structure
├── Timeline
├── Composition
└── Style
```

Every higher layer depends on lower layers but never modifies them.

---

# Layer 1 — Metadata

Metadata describes the recording itself.

Examples

- title
- artist
- album
- duration
- sample rate
- channels
- filename

Metadata is never used directly for generation.

---

# Layer 2 — Audio

Audio is the original waveform.

Usually

```
WAV

or

FLAC
```

The waveform is immutable.

It is used only by analyzers.

---

# Layer 3 — MusicDNA

MusicDNA describes

"What does this music sound like?"

Examples

```
tempo

key

spectral centroid

dynamic range

rhythm density

tonal stability
```

MusicDNA contains global characteristics.

MusicDNA does not know where musical events happen.

---

# Layer 4 — Structure

Structure answers

"What parts does this composition contain?"

Typical sections

```
Intro

Theme

Development

Bridge

Break

Climax

Outro
```

Each section has

- start
- end
- duration
- confidence

---

# Layer 5 — Timeline

Timeline answers

"What happens over time?"

Timeline consists of ordered musical events.

Example

```
00:00

Pad starts

00:18

Noise layer appears

00:42

Bass enters

01:05

Melody begins

02:30

Harmony changes

03:40

Break

04:05

Theme returns

05:10

Outro
```

Timeline is deterministic.

The same song always produces the same timeline.

---

# Musical Event

Every timeline entry is represented as

```
Event

time

duration

type

confidence

parameters
```

Example

```
BassEnter

time = 42.3

confidence = 0.96
```

---

# Event Categories

Events belong to different categories.

Examples

```
Harmony

Rhythm

Texture

Dynamics

Energy

Transition

Instrument

Effect
```

The list may grow over time.

---

# Layer 6 — Composition

Composition answers

"How does the musical idea evolve?"

Unlike Timeline,

Composition is semantic.

Example

```
Atmosphere

↓

Expectation

↓

Development

↓

Release

↓

Resolution
```

This layer describes musical intention rather than acoustic events.

---

# Composition Graph

The composition is represented as a directed graph.

Example

```
Intro

↓

Atmosphere

↓

Main Theme

↓

Development

↓

Break

↓

Return

↓

Ending
```

Nodes represent ideas.

Edges represent transitions.

This graph becomes the primary learning object for future generators.

---

# Layer 7 — Style

Style is extracted from the complete library.

It is never computed from a single track.

Examples

```
preferred BPM

preferred dynamics

preferred harmonic movement

preferred structures

preferred energy evolution

preferred transitions
```

Style describes the author's musical language.

---

# Why Multiple Layers?

Different AI components require different information.

MusicDNA

↓

describes sound

Structure

↓

describes organization

Timeline

↓

describes events

Composition

↓

describes ideas

Style

↓

describes habits

---

# Generator Input

The future Composer never receives raw audio.

Instead it receives

```
Style

+

Composition Graph

+

Timeline Templates

+

MusicDNA Statistics

+

Random Seed
```

---

# Generator Output

The generator produces

```
Composition Graph

↓

Timeline

↓

Musical Events

↓

Rendering

↓

Audio
```

Generation is therefore hierarchical.

---

# Learning Strategy

The system learns progressively.

Stage 1

Understand sound.

Stage 2

Understand structure.

Stage 3

Understand musical ideas.

Stage 4

Understand author style.

Stage 5

Compose new music.

---

# Design Principles

Raw audio is immutable.

MusicDNA is immutable.

Timeline is deterministic.

Composition is semantic.

Style is extracted from the complete library.

Generation consumes abstractions.

Rendering is independent from composition.

---

# Long-Term Goal

AI MusiMuse should eventually understand not only

"What sounds are present?"

but also

"Why does this composition evolve the way it does?"

The generated music should therefore resemble the author's musical thinking rather than copying existing tracks.