# AI MusiMuse

> Master Specification (SPEC)

Version: 1.0

Status: Draft

Document Type: Project Constitution

---

# Project Rule Zero

Whenever there is uncertainty, choose the solution that makes the project easier to understand, easier to extend and easier to explain.

AI MusiMuse is a long-term research project.

Short-term convenience must never compromise long-term architecture.

---

# Table of Contents

1. Vision
2. Philosophy
3. Project Goals
4. Non Goals
5. Design Principles
6. AI Rules
7. Engineering Standards
8. High-Level Architecture
9. Core Modules
10. Data Flow
11. Data Model
12. Measurement Registry
13. Music DNA
14. Confidence Model
15. Versioning
16. Storage Principles
17. Architecture Contracts
18. Plugin SDK
19. Configuration
20. Logging
21. Testing
22. Roadmap

---

# 1. Vision

## 1.1 Purpose

AI MusiMuse is an offline-first research platform whose purpose is to understand the musical identity of a composer.

The project is not intended to replace human creativity.

Instead, it attempts to model musical language, creative habits, stylistic evolution and latent artistic space.

AI MusiMuse should eventually become an intelligent musical companion rather than a music generator.

---

## 1.2 Fundamental Question

Traditional music AI asks:

> Generate a song.

AI MusiMuse asks:

> What musical ideas naturally belong to this composer's artistic universe?

Every architectural decision should support this philosophy.

---

## 1.3 Mission

The system exists to

- understand
- remember
- compare
- explain
- inspire

Generation is only a consequence of understanding.

Understanding is the primary objective.

---

# 2. Philosophy

## Human First

The composer always makes artistic decisions.

The AI assists.

The AI never replaces artistic intent.

---

## Explainability

Every recommendation must be explainable.

Every similarity score must expose measurable evidence.

Every conclusion should reference observable musical properties.

---

## Scientific Thinking

The project should prefer measurable observations over subjective assumptions.

Whenever uncertainty exists it should be represented explicitly.

---

## Offline First

Core functionality must work without Internet access.

Cloud services are optional extensions.

---

## Deterministic Analysis

Running the same analysis twice with identical inputs must produce identical results.

Randomness during analysis is forbidden.

---

## Long-Term Stability

The architecture should remain maintainable for many years.

Temporary convenience must never compromise long-term quality.

---

# 3. Project Goals

The system should eventually be capable of

• analysing complete music libraries

• detecting stylistic evolution

• modelling artistic identity

• discovering recurring compositional habits

• identifying unexplored creative directions

• explaining similarities between tracks

• generating reports

• answering questions about the music library

• assisting composition

• optionally controlling external generation models

The primary goal is understanding.

Generation is secondary.

---

# 4. Non Goals

AI MusiMuse is NOT

- a DAW
- a synthesizer
- an audio editor
- a sampler
- a Suno clone
- a MusicGen clone
- a replacement for creative work

---

# 5. Design Principles

## Simplicity

Simple systems are easier to improve.

Avoid unnecessary abstraction.

---

## Modularity

Every module should perform one clearly defined responsibility.

---

## Replaceability

Every subsystem should be replaceable independently.

---

## Extensibility

New algorithms should be added without changing existing code.

---

## Version Everything

Every algorithm.

Every model.

Every feature.

Every embedding.

Every database schema.

Every report.

Everything.

---

## Explainability over Accuracy

A slightly less accurate but explainable algorithm is preferable to a highly accurate black box whenever practical.

---

## Configuration over Hardcoding

Behavior should be configurable.

Magic constants are discouraged.

---

## Composition over Inheritance

Favor small composable components.

---

# 6. AI Rules

These rules are mandatory for every AI agent contributing to the repository.

1. Never duplicate DSP logic.

2. Never duplicate business logic.

3. Never hardcode feature names.

4. Never silently ignore exceptions.

5. Never hide uncertainty.

6. Every public function requires documentation.

7. Every public class requires documentation.

8. Every module must have a single responsibility.

9. Every module must be independently testable.

10. Avoid global mutable state.

11. Avoid circular dependencies.

12. Use dependency injection where appropriate.

13. Prefer composition over inheritance.

14. Every algorithm must be deterministic unless explicitly documented otherwise.

15. Every important operation must be logged.

16. Never modify historical analysis results.

17. Never overwrite previous algorithm versions.

18. Every semantic descriptor must reference measurable evidence.

19. Every recommendation must contain an explanation.

20. Every similarity score must expose contributing factors.

21. Every feature requires documentation.

22. Every feature requires versioning.

23. Every feature requires units where applicable.

24. Every feature requires expected value range.

25. Every algorithm should expose confidence whenever uncertainty exists.

26. Core functionality must never require cloud APIs.

27. Prefer readability over cleverness.

28. Prefer explicit code over implicit behavior.

29. Avoid hidden side effects.

30. The architecture described in this specification has priority over implementation convenience.

---

# 7. Engineering Standards

Python

3.12+

Formatting

Black

Imports

isort

Linting

Ruff

Typing

Strict typing

Testing

pytest

Documentation

Google Style Docstrings

Architecture

Hexagonal-inspired modular architecture

Database

SQLAlchemy ORM

Configuration

Pydantic Settings

Logging

structlog

Dependency Management

uv

---

# 8. High-Level Architecture

```

                    Music Library

                           │

                           ▼

                  File Discovery

                           │

                           ▼

                  Audio Decoder

                           │

                           ▼

                  DSP Engine

                           │

                           ▼

                 Feature Extraction

                           │

                           ▼

                  Music DNA Engine

                           │

                           ▼

                  Embedding Engine

                           │

                           ▼

                Similarity Engine

                           │

                           ▼

                 Memory Database

                           │

                           ▼

              Recommendation Engine

                           │

                           ▼

                   LLM Reasoner

                           │

                           ▼

                     User Interface

```

Only the LLM Reasoner may contain probabilistic reasoning.

Everything before it must remain deterministic.

---

# 9. Core Modules

## File Scanner

Responsibilities

- recursive discovery
- duplicate detection
- change detection
- hashing
- metadata extraction

Output

Track Manifest

---

## Audio Decoder

Responsibilities

- loading
- normalization
- resampling
- channel handling

Supported formats

- WAV
- FLAC
- AIFF
- MP3

Future

- AAC
- OGG
- M4A

Output

Audio Buffer

---

## DSP Engine

Responsible only for signal processing.

No AI.

No semantic interpretation.

Produces measurable observations.

Examples

- Tempo
- RMS
- LUFS
- Peak
- Spectral Centroid
- Spectral Bandwidth
- Spectral Contrast
- Spectral Flux
- MFCC
- Chroma
- Harmonic Ratio
- Stereo Width
- Dynamic Range

The DSP engine never decides what these values mean.

It only measures.

---

# 10. Data Model

The data model is the foundation of AI MusiMuse.

Every piece of information produced by the system must be represented as structured, versioned and reproducible data.

The database is considered the long-term memory of the composer.

No algorithm may depend on transient runtime state.

---

# 10.1 Core Principles

The database stores facts.

Algorithms produce facts.

Facts never disappear.

Algorithms evolve.

Facts remain reproducible.

Every stored value must answer three questions:

• what was measured

• how it was measured

• by which algorithm version

---

# 10.2 Core Entities

The system revolves around the following entities.

Project

Track

TrackVersion

AnalysisRun

FeatureDefinition

FeatureValue

Embedding

Similarity

Recommendation

Report

ConfigurationSnapshot

AlgorithmVersion

---

# 11. Project

Represents one musical workspace.

Fields

UUID

Name

Composer

Created At

Description

Preferred Language

Default Analysis Version

Project Notes

The project owns every other entity.

---

# 12. Track

A Track represents one musical composition.

A Track never changes.

If the audio changes, a new TrackVersion is created.

Fields

UUID

Project ID

Title

Original Filename

Relative Path

SHA-256

Import Date

Duration

Channels

Sample Rate

Bit Depth

Codec

File Size

Tags

Notes

---

# 13. TrackVersion

Stores revisions of source material.

Example

Original Mix

↓

Mastered Mix

↓

Extended Mix

↓

Remastered Version

Each version has

Version Number

Import Date

Audio Hash

Parent Version

Description

---

# 14. AnalysisRun

Represents one execution of the complete analysis pipeline.

Every execution receives its own identifier.

Fields

UUID

Track ID

Started At

Finished At

Analysis Version

Configuration Snapshot

Git Commit

Success Flag

Execution Time

Errors

Warnings

One track may have many analysis runs.

Historical analyses are never modified.

---

# 15. FeatureDefinition

A FeatureDefinition describes WHAT is being measured.

It never stores measurement values.

Example

Identifier

Brightness

Description

Estimated perceived spectral brightness.

Category

Music DNA

Units

Normalized

Range

0.0 — 1.0

Dependencies

Spectral Centroid

Bandwidth

Spectral Rolloff

Version

1.0

Documentation

Reference URL

Author

Confidence Strategy

FeatureDefinition is immutable.

Changing the calculation creates a new version.

---

# 16. FeatureValue

A FeatureValue stores one measured value.

Every value belongs to exactly

one Track

one AnalysisRun

one FeatureDefinition

Fields

UUID

Track ID

AnalysisRun ID

FeatureDefinition ID

Value

Confidence

Algorithm Version

Calculated At

Processing Time

Every FeatureValue is immutable.

No updates are allowed.

Corrections create new AnalysisRuns.

---

# 17. Measurement Registry

Every measurable property used anywhere in the project MUST be registered.

The registry acts as the vocabulary of AI MusiMuse.

Nothing may reference an undefined feature.

Every registry entry must define

Identifier

Display Name

Description

Units

Expected Range

Category

Dependencies

Algorithm

Version

Confidence Strategy

Example Values

Documentation

Deprecation Status

---

# 18. Measurement Categories

Measurements belong to categories.

Categories improve discoverability.

## Metadata

Duration

Channels

Sample Rate

Bit Depth

Codec

Bitrate

File Size

---

## Dynamics

Peak

RMS

LUFS

Dynamic Range

Crest Factor

Envelope Statistics

---

## Rhythm

Tempo

Beat Stability

Onset Density

Pulse Strength

Transient Density

Groove Consistency

---

## Spectral

Centroid

Bandwidth

Rolloff

Contrast

Flux

Flatness

MFCC

Spectral Slope

---

## Harmonic

Key

Mode

Key Confidence

Chroma

Harmonic Ratio

Tonal Stability

Pitch Class Distribution

---

## Spatial

Stereo Width

Mid Energy

Side Energy

Estimated Reverb

Spatial Depth

Phase Correlation

---

## Semantic (Music DNA)

Brightness

Warmth

Movement

Dreaminess

Space

Nature

Mystery

Energy

Density

Complexity

Calmness

Tension

Emotional Weight

Atmospheric Richness

These descriptors are NEVER measured directly.

They are inferred from measurable observations.

---

# 19. Music DNA

Music DNA is the semantic layer of AI MusiMuse.

DSP answers

"What exists?"

Music DNA answers

"What does it mean?"

Music DNA never reads audio directly.

It consumes FeatureValues.

Every semantic descriptor must define

Input Features

Normalization Strategy

Aggregation Strategy

Output Range

Confidence Strategy

Documentation

Version

Example

Brightness

Inputs

Spectral Centroid

Bandwidth

Spectral Rolloff

Aggregation

Weighted Normalized Average

Output

0.0 — 1.0

Confidence

Derived from input confidence propagation

---

Dreaminess

Inputs

Brightness

Stereo Width

Reverb Estimate

Transient Density

Spectral Flatness

Movement

Aggregation

Weighted Multi-factor Model

Output

0.0 — 1.0

---

Music DNA descriptors may depend on other Music DNA descriptors.

Circular dependencies are forbidden.

The dependency graph must always remain acyclic.

---

# 20. Confidence Model

Whenever uncertainty exists, algorithms should expose confidence.

Confidence expresses certainty.

It never expresses quality.

Examples

Tempo

63.1 BPM

Confidence

0.99

Estimated Key

D Minor

Confidence

0.81

Dreaminess

0.74

Confidence

0.69

Brightness

0.82

Confidence

0.98

Confidence should be based on measurable evidence whenever possible.

Confidence values must be reproducible.

Confidence algorithms are versioned.

Confidence never replaces explanation.

---

# 21. Versioning Strategy

Everything is versioned.

Analysis Version

Algorithm Version

Feature Version

Embedding Version

Configuration Version

Database Schema Version

Similarity Version

Recommendation Version

Historical reproducibility is mandatory.

Nothing is overwritten.

---

# 22. Storage Principles

Raw audio is immutable.

Analysis results are immutable.

Generated content is immutable.

Caches are disposable.

Derived values may always be regenerated.

The database stores knowledge.

The filesystem stores source material.

These responsibilities must never overlap.

---

# 23. Architecture Contracts

The architecture is built around explicit contracts.

Each module has clearly defined responsibilities, inputs and outputs.

Modules communicate only through typed domain objects.

Modules never communicate through database internals.

Modules never access unrelated modules directly.

Each module should be independently testable.

---

# 23.1 File Scanner

Responsibility

Discover source audio files.

Input

Filesystem

Output

TrackManifest

Must

- detect new files
- detect deleted files
- detect modified files
- calculate SHA-256
- ignore unsupported formats

Must NOT

- decode audio
- calculate features
- write recommendations

---

# 23.2 Audio Decoder

Responsibility

Convert audio files into normalized AudioBuffer objects.

Input

TrackManifest

Output

AudioBuffer

Must

- normalize sample format
- preserve sample rate metadata
- preserve channel layout

Must NOT

- calculate DSP
- calculate embeddings
- access database

---

# 23.3 DSP Engine

Responsibility

Extract objective measurable information.

Input

AudioBuffer

Output

DSPFeatureCollection

Must

- remain deterministic
- produce reproducible values
- expose confidence where appropriate

Must NOT

- infer semantics
- classify mood
- generate recommendations

---

# 23.4 Music DNA Engine

Responsibility

Transform objective DSP observations into semantic descriptors.

Input

DSPFeatureCollection

Output

SemanticFeatureCollection

Must

- use only registered FeatureDefinitions
- expose dependency graph
- calculate confidence

Must NOT

- access audio
- modify DSP values

---

# 23.5 Embedding Engine

Responsibility

Generate vector representations.

Input

AudioBuffer
FeatureCollection

Output

Embedding

Supported Models

CLAP

MusicFM

Future embedding models

Multiple embedding models may coexist.

Embeddings are versioned.

---

# 23.6 Similarity Engine

Responsibility

Compare musical works.

Input

Embeddings

FeatureValues

Output

SimilarityGraph

The engine should support

Cosine Similarity

Euclidean Distance

Hybrid Similarity

Weighted Similarity

Every similarity score must expose contributing factors.

Example

Similarity

0.84

Contributors

Brightness

+18%

Warmth

+14%

Movement

+8%

Tempo

-3%

Reverb

+11%

The explanation is part of the result.

---

# 23.7 Recommendation Engine

Responsibility

Transform knowledge into useful creative suggestions.

Input

SimilarityGraph

Music DNA

Historical Statistics

Output

Recommendations

Examples

Try a slower tempo.

Increase harmonic density.

Reduce transient activity.

Explore darker timbres.

Recommendations are suggestions.

Never instructions.

Every recommendation must contain evidence.

---

# 23.8 Report Engine

Responsibility

Generate human-readable reports.

Supported Formats

Markdown

HTML

PDF

JSON

CSV

Reports must be reproducible.

---

# 24. Plugin SDK

The system is extensible through plugins.

Plugins allow researchers and developers to add new analysis modules.

Plugins must never modify core code.

---

Plugin Types

DSP Plugin

Semantic Plugin

Embedding Plugin

Recommendation Plugin

Import Plugin

Export Plugin

Visualization Plugin

---

Every plugin must declare

Identifier

Version

Author

Dependencies

Input Types

Output Types

Configuration Schema

Documentation

License

---

Plugin Lifecycle

Load

↓

Validate

↓

Register

↓

Initialize

↓

Execute

↓

Dispose

---

Plugins may be disabled.

The core system must remain operational.

---

# 25. Configuration Philosophy

Every configurable parameter belongs to a configuration object.

No magic constants are allowed inside algorithms.

Configuration files should be human readable.

Preferred format

YAML

Environment variables should only be used for

Paths

Logging

Development settings

Never for algorithm tuning.

Algorithm parameters belong inside project configuration.

Configuration is versioned.

Configuration snapshots are stored with every AnalysisRun.

---

# 26. Logging

Logging exists for reproducibility.

Every important operation should be logged.

Required Events

Project Opened

Track Imported

Analysis Started

Analysis Finished

Plugin Loaded

Plugin Failed

Embedding Generated

Recommendation Generated

Database Migration

Unexpected Error

Logging Levels

DEBUG

INFO

WARNING

ERROR

CRITICAL

Logs should never contain raw audio.

---

# 27. Error Handling

Errors should never be ignored.

Recoverable errors should be reported.

Unrecoverable errors should stop the current operation.

The system must always explain

what happened

why it happened

how to recover

Silent failures are forbidden.

---

# 28. Testing Strategy

Every module must have automated tests.

Test Types

Unit Tests

Integration Tests

Regression Tests

Golden Tests

Performance Tests

Golden Tests are especially important.

Given identical inputs,

analysis results must remain identical.

Regression testing is mandatory for

DSP

Music DNA

Similarity

Embeddings

---

# 29. Performance Principles

Correctness is more important than speed.

Reproducibility is more important than optimization.

Optimization should be driven by profiling.

Premature optimization is discouraged.

Heavy computations should support caching.

Parallel processing should never affect determinism.

---

# 30. Security

The project is offline-first.

Privacy is a design goal.

No user data should leave the local machine unless explicitly requested.

Generated reports belong to the user.

Source audio always remains under user control.

No telemetry is required.

Anonymous usage statistics are optional and disabled by default.

---

# 31. Repository Layout

The repository should remain clean, predictable and scalable.

Recommended structure:

AI-MusiMuse/

docs/
    SPEC.md
    ARCHITECTURE.md
    ROADMAP.md
    CHANGELOG.md

src/
    core/
    scanner/
    decoder/
    dsp/
    music_dna/
    embeddings/
    similarity/
    recommendations/
    reports/
    plugins/
    database/
    config/
    logging/
    utils/

tests/
    unit/
    integration/
    regression/
    golden/

scripts/

examples/

assets/

output/

cache/

data/
    music/

migrations/

---

# 32. Dependency Rules

Dependencies always point inward.

Example

User Interface
        ↓

Recommendation Engine
        ↓

Similarity Engine
        ↓

Music DNA
        ↓

DSP
        ↓

Decoder
        ↓

Filesystem

Lower layers must never depend on upper layers.

The DSP Engine must never import Recommendation modules.

The Music DNA Engine must never import UI code.

Database entities must never depend on GUI classes.

Circular dependencies are forbidden.

---

# 33. Development Workflow

Development follows an incremental workflow.

Every feature passes through the following stages.

Idea

↓

Specification

↓

Implementation

↓

Tests

↓

Documentation

↓

Review

↓

Merge

Code should never be written before its behavior is understood.

Whenever architecture changes, SPEC.md must be updated before implementation.

---

# 34. Definition of Done

A feature is considered complete only if

✓ implementation finished

✓ documented

✓ tested

✓ deterministic

✓ logged

✓ configurable

✓ versioned

✓ benchmarked (when applicable)

✓ integrated into the pipeline

✓ accepted by regression tests

Incomplete work should never be merged into the main branch.

---

# 35. Release Strategy

Development versions

0.x.x

Public beta

1.0.0-beta

Stable release

1.0.0

Semantic Versioning must be followed.

Major

Breaking architectural changes

Minor

New functionality

Patch

Bug fixes

---

# 36. Roadmap

## Version 0.1

Project bootstrap

Repository

Configuration

Logging

SQLite

Scanner

Track import

---

## Version 0.2

Audio decoding

DSP Engine

Metadata extraction

Golden tests

---

## Version 0.3

Feature Registry

FeatureValue storage

Music DNA v1

Analysis pipeline

---

## Version 0.4

Embeddings

Similarity Engine

Similarity explanations

---

## Version 0.5

Recommendation Engine

Reports

Statistics

Trend analysis

---

## Version 0.6

Plugin SDK

Plugin loader

Plugin registry

---

## Version 0.7

LLM reasoning

Natural language queries

Composer assistant

---

## Version 0.8

Creative suggestions

Style evolution

Idea exploration

---

## Version 0.9

Performance optimization

Large library support

Parallel execution

Caching improvements

---

## Version 1.0

Production-ready release

Complete documentation

Stable APIs

Migration tools

Long-term support

---

# 37. Future Research

The following ideas are intentionally outside Version 1.0.

Automatic motif detection

Harmony graph analysis

Arrangement analysis

Structural segmentation

Instrument recognition

Synthesizer fingerprinting

Reverb fingerprint estimation

Mixing style analysis

Composer evolution timeline

Creative anomaly detection

Style interpolation

Latent style navigation

Prompt generation for external music models

Cross-project comparison

Knowledge graph visualization

These features should remain optional.

The core architecture must not depend on them.

---

# 38. Guiding Principles

When making engineering decisions, prefer

clarity over cleverness

simplicity over complexity

knowledge over assumptions

determinism over randomness

measurement over intuition

architecture over shortcuts

understanding over generation

---

# 39. Project Motto

Understand.

Measure.

Remember.

Reason.

Inspire.

Generation is only the consequence of understanding.

---

# 40. Final Statement

AI MusiMuse is not intended to imitate a composer.

It is intended to understand a composer's musical language.

The purpose of the project is to build a long-term musical knowledge system capable of explaining artistic identity, tracking creative evolution and assisting future composition.

Every architectural decision should support this vision.

Whenever uncertainty exists, follow Project Rule Zero.
