# TASK-002

## Title

File Scanner and Track Import

---

## Objective

Implement the first music library processing stage according to **SPEC.md**.

This task introduces filesystem scanning, immutable track discovery and SQLite persistence.

The implementation should establish the foundation for every future analysis pipeline.

Do **NOT** implement audio decoding, DSP, Music DNA, embeddings, similarity, recommendations or any analysis logic.

---

# Expected Commit Message

Add file scanner and track import

---

# Requirements

Before writing any code:

* Read **SPEC.md**
* Read **AI_DEVELOPER_GUIDE.md**
* Read **PROJECT_PRINCIPLES.md**

Follow all project rules.

If architectural uncertainty exists, stop and ask.

Do not guess.

---

# Architecture

Responsibilities must remain clearly separated.

```
Filesystem
      ↓
Scanner
      ↓
TrackManifestCollection
      ↓
Importer
      ↓
Repository
      ↓
SQLite
```

The Scanner discovers files.

The Importer synchronizes the database.

Repositories perform persistence.

These responsibilities must never overlap.

---

# Scope

Implement the following modules.

## Scanner

Responsibilities

* recursively scan the configured music directory
* detect supported audio files
* calculate SHA-256 hashes
* identify duplicate files
* detect new files
* detect modified files
* detect deleted files
* ignore unsupported files

The Scanner must not access SQLite.

The Scanner must not import SQLAlchemy.

The Scanner must only inspect the filesystem.

---

## Track Manifest

Create immutable domain objects representing the current filesystem state.

Required objects

* TrackManifest
* TrackManifestCollection

TrackManifest must contain only filesystem information.

It must not depend on SQLAlchemy models.

It must not contain database logic.

---

## Importer

Responsibilities

* synchronize TrackManifestCollection with the database
* create a default project if none exists
* import new tracks
* detect deleted tracks
* detect modified tracks
* detect duplicate hashes

The Importer must never scan directories directly.

The Importer must consume TrackManifestCollection only.

---

## Repository Layer

Create

ProjectRepository

TrackRepository

Responsibilities

ProjectRepository

* get default project
* create default project

TrackRepository

* create track
* find by path
* find by SHA-256
* list tracks
* mark track as missing

Business logic must not execute raw SQL directly.

Only repositories should communicate with SQLAlchemy.

---

# Database Models

Create SQLAlchemy models.

## Project

Required fields

* UUID
* Name
* Created At

---

## Track

Required fields

* UUID
* Project ID
* Relative Path
* Original Filename
* SHA-256
* Import Date
* File Size
* Status

Audio metadata remains nullable.

Do not decode audio.

The following fields must exist but remain NULL until future tasks:

* Duration
* Channels
* Sample Rate
* Bit Depth

---

# File Hashing

SHA-256 must be calculated using streamed reads.

Never load the entire file into memory.

---

# Supported Formats

Support only

* wav
* flac
* aiff
* aif

Ignore every other extension.

---

# Immutable History

Tracks are immutable.

If a file changes while keeping the same path:

* do not overwrite the existing Track
* do not update the existing SHA-256
* report the modification

TrackVersion will be introduced in a future task.

---

# Deleted Files

Deleted tracks must never be removed from the database.

Mark them as missing.

History must remain intact.

---

# Forbidden

Do NOT

* decode audio
* read metadata from audio
* calculate DSP features
* calculate Music DNA
* create embeddings
* implement similarity
* implement recommendations
* implement TrackVersion
* implement analysis pipeline

---

# Logging

Log important events.

Examples

* scan started
* scan finished
* track imported
* duplicate detected
* deleted track detected
* modified track detected

Do not log excessive details.

---

# Testing

Required tests

* recursive scanning
* supported formats
* unsupported formats
* SHA-256 calculation
* duplicate detection
* deleted file detection
* modified file detection
* default project creation
* repository operations
* importer synchronization

All tests must be deterministic.

---

# Quality Requirements

All code must

* use type hints
* follow PEP-8
* contain Google-style docstrings
* pass Ruff
* pass Black
* pass isort
* pass pytest

---

# Deliverables

At the end of this task the project must

* discover music files in `data/music`
* create a default project automatically
* import new tracks
* detect deleted files
* detect modified files
* detect duplicate files
* ignore unsupported files
* synchronize the SQLite database

No audio analysis should exist.

---

# Definition of Done

The task is complete only if

✓ recursive scanning works

✓ supported formats are discovered

✓ unsupported formats are ignored

✓ SHA-256 hashes are calculated correctly

✓ duplicate files are detected

✓ TrackManifestCollection is produced

✓ repositories are implemented

✓ a default project is created automatically

✓ new tracks are stored in SQLite

✓ deleted tracks are marked as missing

✓ modified tracks are reported without overwriting immutable records

✓ tests pass

✓ Ruff passes

✓ Black passes

✓ isort passes

Nothing else should be implemented.
