Files
dl-organizer/README.md
T
6f0df5a774 release: v0.2.0 repository hygiene, CI, and docs sync
Stop tracking personal workflow artifacts at repo root, add CI and MIT
license, align README and agent skills with artifacts/ defaults, and
enable Ruff in dev/CI so releases are verifiable without local-only runs.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-01 15:32:03 +08:00

941 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Video Library Manager
## Documentation Status
- Last synchronized: **2026-06-01** (see [`CHANGELOG.md`](CHANGELOG.md) for release history)
- Validation baseline: **`uv run pytest -q` → 517 passed**
- Human-in-the-loop workflow includes `vlm apply-review` for syncing manual plan edits from CSV back into `plan.json`.
- Scanner now detects `ffprobe` availability, degrades gracefully, and treats partial `find` output as a warning-backed partial result.
- JSON artifacts (`identities.json`, `analysis.json`, `plan.json`) are schema-validated on load/save, and plan loading now crosses a typed `ExecutionPlan` boundary.
- Plan generation separates logical intent from environment-derived validation snapshots, improving reproducibility.
- `vlm review-plan --tui` remains an optional Textual dependency; the default CLI path does not require Textual.
- Execution rejects move/rename plans whose source or destination escapes `library_root`.
A Python-based CLI tool for managing personal video collections with a safety-first, human-in-the-loop approach.
## Features
- **Safety-First Design**: All file operations are reversible with rollback support
- **Human-in-the-Loop**: Explicit confirmation required before making any changes. **New: Full review-apply cycle for manual plan adjustments via CSV.**
- **Comprehensive Analysis**: Detect episode gaps and duplicate files
- **Rich Metadata**: Extract video resolution, codec, duration, and bitrate (with graceful fallback if `ffprobe` is missing)
- **Flexible Organization**: Customizable directory structure and naming templates
- **Metadata Enrichment**: Add bilingual titles and reputation signals (TMDB + optional AI fallback)
- **Incremental Performance**: SQLite-backed cache avoids repeated metadata lookups
- **State Tracking**: Track file status throughout the workflow
- **Detailed Reporting**: Generate inventory, completeness, and duplicate reports
- **PlanAnalysis Integration**: `vlm plan --analysis` applies duplicate resolution and adds a Chinese human summary
- **Artifact Validation**: JSON artifacts are validated early to catch malformed inputs before later stages run
- **Deterministic Planning**: Plan generation records live filesystem checks as validation snapshots instead of mixing them into core plan facts
## Installation
This project uses `uv` for Python package management. To install:
```bash
# Install dependencies
uv pip install -e .
# Install with development dependencies
uv pip install -e ".[dev]"
# Optional: Textual TUI for `vlm review-plan --tui` (not included in [dev])
uv pip install -e ".[tui]"
# Or both: uv pip install -e ".[dev,tui]"
```
The standard CLI remains fully usable without Textual; the dependency is imported only when `vlm review-plan --tui` is requested.
## Quick Start
### 1. Initialize Configuration
First, create a configuration file:
```bash
vlm config init
```
This creates `~/.vlm/config.yaml`. Edit it to set your library root:
```yaml
library_root: "/mnt/Downloads" # Change this to your video library path
workspace_dir: "artifacts" # Default workspace for generated files
```
### 2. Scan Your Library
Discover all video files in your library:
```bash
vlm scan
```
This creates `artifacts/inventory.csv` with all discovered files and their metadata.
### Generated artifacts (do not commit)
Default pipeline outputs live under `artifacts/` (or `workspace_dir` from config). **Do not commit** `artifacts/`, `runs/`, or legacy root-level files such as `inventory.csv`, `identities.json`, `analysis.json`, or `plan.json` at the repository root. Keep API keys in `~/.vlm/config.yaml` only — never commit config copies or paste keys into issues.
**What happens:**
- Discovers video files using the system `find` command
- If `find` exits non-zero after emitting paths, VLM keeps the partial results and logs a warning; if no paths are emitted, discovery returns an empty deterministic result
- Extracts file metadata (size, modification time)
- Categorizes files based on directory structure (movie/series/anime/other)
- Extracts video metadata using ffprobe (if available)
- Saves results to `artifacts/inventory.csv`
### 3. Parse Filenames
Extract titles, years, seasons, and episodes from filenames:
```bash
vlm parse
```
**For accurate duplicate resolution by quality, embed video metadata:**
```bash
vlm parse --inventory artifacts/inventory.csv
```
This creates `artifacts/identities.json` with parsed information. When using `--inventory`, video metadata (size, resolution, codec) is embedded, enabling accurate quality comparison during duplicate analysis.
**What it extracts:**
- **Movies**: Title and year (e.g., "Inception (2010)")
- **Series**: Title, season, and episode numbers (e.g., "Breaking Bad S01E01")
- **Confidence scores**: Indicates parsing reliability
- **Video metadata** (with `--inventory`): Size, resolution, codec, duration, bitrate
### 4. Enrich Titles and Reputation (Optional but Recommended)
Add translation and reputation metadata to `artifacts/identities.json`:
```bash
vlm enrich
```
This updates `artifacts/identities.json` in place and adds fields like:
- `title_zh`, `title_en`, `display_title`
- `reputation_score`, `reputation_votes`, `reputation_source`
- `review_status`, `enrichment_confidence`
- summary metrics including `api_calls`, `cache_hits`, and `skip_reasons`
To refresh all records instead of using incremental cache:
```bash
vlm enrich --refresh-all
```
If enrichment cannot run for some records, CLI shows grouped reasons, for example:
```text
Skip reasons: no_key=4632
```
### 5. Analyze Your Library
Detect episode gaps and duplicates:
```bash
vlm analyze
```
This creates `artifacts/analysis.json` with:
- Series with missing episodes
- Duplicate files with quality comparison
### 6. Generate Execution Plan
Create a reviewable plan of file operations:
```bash
vlm plan
```
To let the plan automatically resolve duplicate groups (keep one file per group by reputation, quarantine the rest), pass the analysis file:
```bash
vlm plan --analysis artifacts/analysis.json
```
This creates `artifacts/plan.json` with:
- Proposed operations (move, rename, quarantine, no-op)
- **Summary**: counts by operation type and by reason
- **Human summary** (中文): short narrative for quick review
- **Metadata**: when using `--analysis`, duplicate groups considered and completeness gaps
Safety guards in planning:
- Titles used in path templates are sanitized (path separators/control chars/`..` stripped)
- Any destination outside `library_root` is rejected as `no-op`
- Duplicate resolution with `--analysis` is deterministic: unresolved `by_quality` groups are converted into explicit manual-review `no-op` operations instead of silently keeping the first file
Duplicate keep strategy is configurable in `~/.vlm/config.yaml` under `plan.duplicate_keep`:
- `by_quality` - Prefer highest quality (resolution > source > codec > file size). Best for automatic duplicate resolution.
- `by_reputation` - Prefer external rating (TMDB); when rating ties/missing, fallback to quality (source-first: BluRay > WEB-DL), then first-seen. Default.
- `by_reputation_quality_time` - Prefer external rating first, then quality, then newer modified time.
- `first_seen` - Keep the first file in each duplicate group.
- `manual` - Do not generate quarantine operations; duplicates are listed in analysis only.
**Review the plan** in one of four ways:
- Open `artifacts/plan.json` in your editor
- Run `vlm review-plan` to get a terminal preview (summary + high-risk operation preview)
- Run `vlm review-plan --tui` for the optional full-screen Textual review UI
- Run `vlm execute` to see the same plan summary in dry-run mode
You can still edit `plan.json` directly when needed.
### 7. Execute (Dry-Run First)
Preview what will happen without making changes:
```bash
vlm execute
```
Before running, the CLI prints the plans **human summary** (or a short summary from counts) so you can confirm at a glance. When ready to actually move/rename/quarantine files:
```bash
vlm execute --confirm
```
**Important**: This creates a rollback log in `~/.vlm/rollback/` for reverting changes.
Execution safeguards:
- Even if a manually edited `artifacts/plan.json` contains an unsafe move/rename source or destination, execution rejects paths outside `library_root`
- Unsupported quarantine operations are reported as failed results without aborting later operations in the same run
- Summary counters treat conflict skips separately from real failures (`failed`/`skipped` are mutually exclusive)
### 8. Rollback (If Needed)
If you need to undo the operations:
```bash
vlm rollback
```
This uses the most recent rollback log to restore files to their original locations.
## Complete Workflow Example
Here's a complete workflow from start to finish:
```bash
# 1. Initialize configuration
vlm config init
# Edit ~/.vlm/config.yaml to set your library_root
# 2. Scan your library
vlm scan
# Output: artifacts/inventory.csv with 1234 files discovered
# 3. Parse filenames with metadata embedding (recommended for duplicate resolution)
vlm parse --inventory artifacts/inventory.csv
# Output: artifacts/identities.json with parsed titles, episodes, and embedded video metadata (v2 schema)
# 4. Enrich identities (translation + reputation)
vlm enrich
# Output: artifacts/identities.json updated in place (incremental cache enabled)
# 5. Analyze for gaps and duplicates
vlm analyze
# Output: artifacts/analysis.json with 5 series with gaps, 12 duplicate groups (accurate quality comparison)
# 6. Generate execution plan (optionally use analysis for duplicate handling)
vlm plan --analysis artifacts/analysis.json
# Output: artifacts/plan.json with operations, human summary, and duplicate quarantine decisions
# 7. Review the plan in terminal (summary + high-risk preview + verdict)
vlm review-plan
# Optional: interactive full-screen review (install `.[tui]` first)
vlm review-plan --tui
# Enriched CSV (title, relative paths, quality) + duplicate grouping
vlm review-plan --identities artifacts/identities.json --analysis artifacts/analysis.json --group-by duplicate
# Spot-check random safe moves; write target tree preview
vlm review-plan --sample-safe 5 --structure-preview artifacts/plan_structure.txt
# Edit artifacts/plan_manual_review.csv in Excel/Numbers
# Sync your manual decisions back to artifacts/plan.json
vlm apply-review
# Output: Successfully updated plan saved to: artifacts/plan.json
# 8. Dry-run to preview
vlm execute
# Shows what would happen without making changes (respecting your manual edits)
# 9. Execute with confirmation (optional review gate)
vlm execute --confirm --require-review
# Actually performs the file operations
# 10. If needed, rollback
vlm rollback
```
## Command Reference
### Configuration
```bash
# Initialize config file
vlm config init
# Show current configuration
vlm config show
# Validate configuration
vlm config validate
```
### Scanning
```bash
# Scan with default output (artifacts/inventory.csv)
vlm scan
# Scan with custom output file
vlm scan --output my_library.csv
```
**Scan behavior:**
- Hidden paths are skipped (any component starting with `.`)
- File matching uses configured `video_extensions` (default: `.mp4`, `.mkv`, `.avi`, `.mov`, `.wmv`, `.flv`, `.webm`, `.m4v`)
- If `ffprobe` is available, resolution/codec/duration/bitrate are included
- ffprobe extraction uses controlled concurrency (`enrichment.max_concurrency`)
- Timestamps stored in UTC format (`YYYY-MM-DDTHH:MM:SS`)
### Parsing
```bash
# Parse with default files (v1 schema - no metadata embedding)
vlm parse
# Parse with metadata embedding (v2 schema - enables quality comparison)
vlm parse --inventory artifacts/inventory.csv
# Parse with custom input/output
vlm parse --input my_inventory.csv --output my_identities.json
# Parse with metadata from custom inventory
vlm parse --input my_inventory.csv --output my_identities.json --inventory my_inventory.csv
```
**Schema Versions:**
- **v1** (without `--inventory`): Lightweight identities, suitable for basic organization
- **v2** (with `--inventory`): Embeds video metadata, required for accurate duplicate resolution by quality
### Enrichment
```bash
# Enrich identities in place (default: artifacts/identities.json)
vlm enrich
# Enrich custom file and write to another file
vlm enrich --input my_identities.json --output enriched_identities.json
# Refresh changed records only (explicit incremental mode)
vlm enrich --refresh-changed-only
# Force full refresh (ignore cache for all records)
vlm enrich --refresh-all
# Tune request behavior
vlm enrich --timeout 6 --retries 2
```
### Analysis
```bash
# Analyze with default files
vlm analyze
# Analyze with custom files
vlm analyze --input my_identities.json --output my_analysis.json
```
### Planning
```bash
# Generate plan with default files
vlm plan
# Use analysis so duplicate groups become "keep one + quarantine rest" (by_reputation by default)
vlm plan --analysis artifacts/analysis.json
# Custom input/output
vlm plan --input my_identities.json --output my_plan.json
vlm plan --input my_identities.json --analysis my_analysis.json --output my_plan.json
```
### Manual Plan Review
```bash
# Export high-risk operations and preview them in terminal
vlm review-plan
# Interactive Textual UI (install optional dependency: uv pip install -e ".[tui]")
vlm review-plan --tui
# Preview first N high-risk operations in terminal (default: 10)
vlm review-plan --preview-limit 20
# Show all high-risk operations in terminal preview
vlm review-plan --show-all
# Apply edited CSV decisions back to plan.json
vlm apply-review
```
### Execution
```bash
# Dry-run (preview only, no changes)
vlm execute
# Execute with confirmation (actually makes changes)
vlm execute --confirm
# Use custom plan file
vlm execute --plan my_plan.json --confirm
```
### Rollback
```bash
# Rollback using most recent log
vlm rollback
# Rollback using specific log
vlm rollback --log ~/.vlm/rollback/rollback_<uuid>.json
```
### Quarantine Management
```bash
# List quarantined files
vlm quarantine list
# List by category
vlm quarantine list --category movie
# Quarantine a file
vlm quarantine add /path/to/file.mkv --reason "duplicate"
# Restore from quarantine
vlm quarantine restore /path/to/.quarantine/file.mkv
```
### Reporting
```bash
# Generate inventory report (text format)
vlm report inventory
# Generate in different formats
vlm report inventory --format csv
vlm report inventory --format json --output inventory_report.json
# Generate completeness report (series with gaps)
vlm report completeness
# Include plan content summary in the report (human_summary from artifacts/plan.json)
vlm report completeness --plan artifacts/plan.json
# Generate duplicates report
vlm report duplicates
vlm report duplicates --plan artifacts/plan.json
# Generate summary statistics
vlm report summary
```
### State Management
```bash
# Show state for a file
vlm state show /path/to/file.mkv
# Set state for a file
vlm state set /path/to/file.mkv --status reviewed
vlm state set /path/to/file.mkv --status ignored --reason "duplicate"
# Query files by status
vlm state query --status ignored
# Clear state for a file
vlm state clear /path/to/file.mkv
```
## Configuration File
The configuration file (`~/.vlm/config.yaml`) controls VLM's behavior:
```yaml
# Required: Root directory of your video library
library_root: "/mnt/nas/videos"
# Workspace directory for generated artifacts
workspace_dir: "artifacts"
# Video file extensions to recognize
video_extensions:
- .mp4
- .mkv
- .avi
- .mov
- .wmv
- .flv
- .webm
- .m4v
# Directory structure templates
templates:
movie_dir: "movie/{title} ({year})/"
series_dir: "series/{title}/Season {season:02d}/"
movie_filename: "{title} ({year}){ext}"
series_filename: "S{season:02d}E{episode:02d}{ext}"
# Quarantine directory name (relative to category root)
quarantine_dir: ".quarantine"
# Logging level (DEBUG, INFO, WARNING, ERROR)
log_level: "INFO"
# Plan behavior (e.g. when using vlm plan --analysis)
plan:
# Duplicate keep strategy: "by_quality", "by_reputation" (default),
# "by_reputation_quality_time", "first_seen", or "manual"
# by_quality: resolution > source (BluRay > WEB-DL) > codec (x265 > x264) > file size
duplicate_keep: "by_reputation"
# Category mappings (directory name to category)
categories:
movie: [movie, movies, films]
series: [series, tv, shows]
anime: [anime]
# Enrichment settings (`enrich` alias is also supported)
enrichment:
enabled: true
incremental: true
refresh_mode: "manual"
providers: [tmdb]
cache_db: "~/.vlm/enrichment_cache.db"
max_concurrency: 6 # used by enrichment requests and scan-time ffprobe workers
min_match_score: 0.75
translation:
mode: "bidirectional"
fallback_machine: true
api_keys:
# Preferred: TMDB v4 Bearer token
tmdb_bearer: null
# Backward-compatible fallback (legacy query api_key)
tmdb: null
openai: null
tmdb:
language: "zh-CN"
region: null
include_adult: false
reputation:
min_votes: 50
low_score_threshold: 6.0
policy: "flag_for_review"
naming:
title_format: "{title_zh} {title_en}"
```
### TMDB Enrichment Setup
`vlm enrich` works best with TMDB Bearer auth (recommended by TMDB). Legacy `tmdb` api key is still supported for compatibility.
Minimal config:
```yaml
enrichment:
providers: [tmdb]
api_keys:
tmdb_bearer: "YOUR_TMDB_BEARER_TOKEN"
```
Optional TMDB query tuning:
```yaml
enrichment:
tmdb:
language: "zh-CN" # localized title language
region: "US" # affects regional release/search behavior
include_adult: false
```
Validation flow:
```bash
# 1) run small incremental pass
vlm enrich --input artifacts/identities.json --refresh-changed-only
# 2) then full refresh if output looks correct
vlm enrich --input artifacts/identities.json --refresh-all
```
### Template Variables
**Movies:**
- `{title}` - Movie title
- `{year}` - Release year
- `{ext}` - File extension
**Series:**
- `{title}` - Series title
- `{season}` - Season number
- `{episode}` - Episode number
- `{ext}` - File extension
Format specifiers are supported (e.g., `{season:02d}` for zero-padded numbers).
## File Categorization
VLM categorizes files based on the top-level directory in your library, matched against configured category mappings. By default:
- `movie/` → "movie"
- `series/` → "series"
- `anime/` → "anime"
- Anything else → "other"
You can customize these mappings to support multiple directory names per category:
```yaml
categories:
movie: [movie, movies, films]
series: [series, tv, shows, television]
anime: [anime]
```
This allows directories like `/library/movies/` or `/library/tv/` to be correctly categorized. Directory matching is case-insensitive.
Example:
```
/mnt/Downloads/
├── movies/ # Category: movie (plural form)
│ └── Inception (2010).mkv
├── tv/ # Category: series (tv variant)
│ └── Breaking Bad S01E01.mkv
├── anime/
│ └── Attack on Titan E01.mkv # Category: anime
└── documentaries/
└── Planet Earth.mkv # Category: other
```
**Migration Note:** If you have existing directories with non-standard names (like "movies" or "tv"), update your `categories` configuration in `config.yaml` and re-run `vlm scan` to fix categorization. No files will be moved.
## Common Scenarios
### Scenario 1: First Time Setup
```bash
# 1. Install
uv pip install -e .
# 2. Initialize config
vlm config init
# Edit ~/.vlm/config.yaml
# 3. Scan library
vlm scan
# 4. Check what was found
vlm report summary
```
### Scenario 2: Finding Missing Episodes
```bash
# 1. Scan and parse
vlm scan
vlm parse --inventory artifacts/inventory.csv
vlm enrich
# 2. Analyze completeness
vlm analyze
# 3. View report (optionally include plan summary if you have a plan)
vlm report completeness
vlm report completeness --plan artifacts/plan.json
```
### Scenario 3: Finding and Removing Duplicates
```bash
# 1. Scan and parse (with --inventory for accurate quality comparison)
vlm scan
vlm parse --inventory artifacts/inventory.csv
vlm enrich
# 2. Analyze for duplicates
vlm analyze
# 3. View duplicates with quality comparison (now shows actual resolution/codec/size)
vlm report duplicates
# 4. Generate plan with analysis: VLM keeps one file per duplicate group (by reputation) and quarantines the rest
vlm plan --analysis artifacts/analysis.json
# 5. Review plan (human summary in artifacts/plan.json and when you run execute)
vlm execute
vlm execute --confirm
# Alternatively: manual quarantine without plan
vlm quarantine add /path/to/lower/quality/file.mkv --reason "duplicate - lower quality"
# Report with plan context
vlm report duplicates --plan artifacts/plan.json
```
### Scenario 4: Reorganizing Your Library
```bash
# 1. Customize templates in config
# Edit ~/.vlm/config.yaml to set your preferred structure
# 2. Scan and parse
vlm scan
vlm parse --inventory artifacts/inventory.csv
vlm enrich
# 3. Generate plan
vlm plan
# 4. Review artifacts/plan.json carefully
# 5. Dry-run to preview
vlm execute
# 6. Execute when satisfied
vlm execute --confirm
# 7. If something goes wrong, rollback
vlm rollback
```
## Logging
VLM logs to `~/.vlm/vlm.log` by default. If the log directory is unwritable, it falls back to console logging and continues running.
## Troubleshooting
### ffprobe Not Found
If you see warnings about ffprobe:
```bash
# Install ffmpeg (includes ffprobe)
# Ubuntu/Debian
sudo apt install ffmpeg
# macOS
brew install ffmpeg
# Windows
# Download from https://ffmpeg.org/download.html
```
Without ffprobe, VLM still works but won't extract video metadata (resolution, codec, etc.).
### Configuration Errors
If VLM can't find your config:
```bash
# Check config location
ls -la ~/.vlm/config.yaml
# Reinitialize if needed
vlm config init
# Validate current config
vlm config validate
```
### TMDB Auth / Rate Limit / Zero Enriched
Common enrichment outcomes:
- `Error during enrichment: TMDB authentication failed (401/403)`
Cause: invalid/missing `tmdb_bearer` (or `tmdb`) key.
Action: update `~/.vlm/config.yaml` and rerun.
- `Skip reasons: no_key=...`
Cause: no TMDB credentials configured for provider.
Action: set `enrichment.api_keys.tmdb_bearer` (recommended) or `tmdb`.
- `Skip reasons: rate_limited=...`
Cause: TMDB rate limit hit (`429`).
Action: retry later; VLM already applies bounded retry/backoff. For large libraries, lower `enrichment.max_concurrency` (try **24** before the default **6**) to reduce simultaneous TMDB requests during `vlm enrich` and ffprobe work during `vlm scan`.
- `Enriched now: 0` with non-zero records
Cause: often `no_key`, `no_match`, or provider errors.
Action: check `Skip reasons` and `Failure sample` in CLI output.
### Permission Errors
If you can't access certain files:
```bash
# Check permissions
ls -la /path/to/library
# VLM logs errors but continues scanning other files
# Check log file for details
cat ~/.vlm/vlm.log
```
### Rollback Fails
Rollback is **best-effort** and may fail if:
- Files have been moved/deleted since execution
- Destination paths are occupied
- Permissions have changed
Always keep backups of important files!
## Safety Features
- **No Permanent Deletion**: VLM never deletes files. Use quarantine instead.
- **Dry-Run Default**: Execute command defaults to dry-run mode
- **Explicit Confirmation**: `--confirm` flag required to make actual changes
- **Rollback Logs**: All operations logged for reversal
- **Conflict Detection**: Plan generation detects destination conflicts
- **Best-Effort Rollback**: Attempt to restore files to original locations
## Limitations
### Version 1.0
- **Anime parsing not implemented**: Anime files are discovered and categorized but not parsed
- **Movie/Series only for quarantine**: Only movie and series categories can be quarantined
- **Hardcoded parsing patterns**: Filename patterns are not user-configurable
- **Local files only**: Designed for local or mounted network storage
- **Best-effort rollback**: Rollback may not succeed if files have been modified
## Agent skills
This project includes workflow guidance under `skills/vlm-library-workflow/` (scan → parse → enrich → analyze → plan → review-plan → apply-review → execute).
### Activation (Gemini CLI)
```bash
activate_skill vlm-library-workflow
```
### Features
- **Context-Aware Guidance**: The Agent understands the VLM safety-first workflow and configuration.
- **Risk Warnings**: Automatic alerts if a plan contains high-risk operations (>20% quarantine).
- **Command Recipes**: Instant access to complex command sequences and troubleshooting steps.
- **Developer Verification**: Built-in test execution recipes for verifying core logic changes.
## Development
### Running Tests
```bash
# Run all tests (recommended)
uv run pytest -q
# Run specific test file
uv run pytest tests/test_scanner.py
# Run with verbose output
uv run pytest -v
# Run with coverage (requires dev extras)
uv run pytest --cov=vlm tests/
```
### Code Style
- Python 3.10+ idioms
- PEP 8 naming conventions
- Type hints for public functions
- 4-space indentation
### Project Structure
```
src/vlm/
├── cli.py # Click-based CLI interface, global options
├── context.py # CLIContext and pass_context for commands
├── cli_helpers.py # Shared CLI helpers
├── commands/ # Command implementations
│ ├── scan.py # Scan command
│ ├── parse.py # Parse command
│ ├── enrich.py # Enrich command
│ ├── analyze.py # Analyze command
│ ├── plan.py # Plan command
│ ├── execute.py # Execute/Rollback commands
│ ├── review_plan.py # Plan review export / TUI
│ ├── report.py # Report commands
│ ├── quarantine_cmd.py
│ ├── state_cmd.py
│ └── config_cmd.py
├── scanner.py # File discovery and metadata extraction
├── parser.py # Filename parsing (titles, years, episodes)
├── enrichment.py # Title/reputation enrichment pipeline
├── cache.py # SQLite cache for incremental enrichment
├── providers/ # External metadata providers (TMDB, etc.)
│ ├── base.py # Provider interface
│ └── tmdb.py # TMDB API client
├── io.py # Unified JSON/CSV I/O helpers (load/save JSON, analysis writer, plan/analysis adapters)
├── utils.py # UTC time, format_size, etc.
├── analysis.py # Completeness and duplicate detection
├── duplicate_resolve.py # Duplicate group resolution strategies
├── planner.py # Execution plan generation (optionally consumes analysis)
├── plan_review.py # Review CSV and apply-review
├── plan_render.py # Plan rendering helpers
├── plan_structure_preview.py
├── review_display.py # Terminal review output
├── review_tui.py # Optional Textual review UI
├── transaction.py # Execution transaction log
├── executor.py # File operations and rollback
├── quarantine.py # Quarantine management
├── state.py # File state tracking
├── reports.py # Report generation
├── config.py # Configuration management
├── models.py # Data structures
└── logging_config.py # Logging setup
tests/
├── test_scanner.py # Scanner tests
├── test_parser.py # Parser tests
├── test_enrichment.py # Enrichment tests
├── test_analysis.py # Analysis tests
├── test_planner.py # Planner tests
└── ... # More test files
```
## Requirements
- **Python** >= 3.10
- **uv** (Python package manager)
- **ffmpeg** (optional, for video metadata extraction)
- **find** command (standard on Unix-like systems)
## Contributing
See [`CONTRIBUTING.md`](CONTRIBUTING.md) for versioning, test commands, and PR expectations.
## License
MIT — see [`LICENSE`](LICENSE).
## Acknowledgments
Built with:
- [Click](https://click.palletsprojects.com/) - CLI framework
- [PyYAML](https://pyyaml.org/) - Configuration parsing
- [pytest](https://pytest.org/) - Testing framework
- [hypothesis](https://hypothesis.readthedocs.io/) - Property-based testing
## Support
- Report issues in your repository host's issue tracker
- Documentation: [`CLAUDE.md`](CLAUDE.md) for architecture; [`CHANGELOG.md`](CHANGELOG.md) for releases
- Logs: Check `~/.vlm/vlm.log` for detailed information