From c9d22d5136a19b9342da08ce984437f7da03233f Mon Sep 17 00:00:00 2001 From: windyboy Date: Tue, 7 Apr 2026 11:07:01 +0800 Subject: [PATCH] docs sync post-refactor baseline --- AGENTS.md | 2 +- ARCHITECTURE_REVIEW.md | 2 +- AUDIT_FIX_PLAN.md | 2 +- CHANGELOG.md | 13 ++++++ CLAUDE.md | 54 +++++++++++++---------- CODE_IMPROVEMENTS.md | 2 +- FIX_PLAN.md | 2 +- GEMINI.md | 14 +++--- IMPLEMENTATION_PLAN_2026-02-13.md | 2 +- IMPROVEMENT_RECOMMENDATIONS_2026-02-13.md | 2 +- README.md | 19 +++++--- TMDB_REFACTOR_PLAN.md | 2 +- VLM_PROJECT_AUDIT_REPORT.md | 2 +- codex_review.md | 2 +- 14 files changed, 75 insertions(+), 45 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 87b2e79..bbb2828 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -39,5 +39,5 @@ ## Documentation baseline -- Updated to reflect refactor results as of 2026-02-16. +- Updated to reflect refactor results as of 2026-04-07. - Canonical release notes are tracked in `CHANGELOG.md`. diff --git a/ARCHITECTURE_REVIEW.md b/ARCHITECTURE_REVIEW.md index ff49fe5..7d42c12 100644 --- a/ARCHITECTURE_REVIEW.md +++ b/ARCHITECTURE_REVIEW.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # Video Library Manager - Architecture Review diff --git a/AUDIT_FIX_PLAN.md b/AUDIT_FIX_PLAN.md index 5d2b086..86b83fe 100644 --- a/AUDIT_FIX_PLAN.md +++ b/AUDIT_FIX_PLAN.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # Post-Audit Fixes Implementation Plan (2026-02-13) diff --git a/CHANGELOG.md b/CHANGELOG.md index 24d107e..f0240aa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,18 @@ # Changelog +## 2026-04-07 + +### Review-plan Safety & Validation Hardening + +- Made `vlm review-plan --tui` a true optional runtime boundary by lazy-loading the Textual UI and guarding Textual imports. +- Hardened execution safety so move/rename operations validate both source and destination against `library_root`, while unsupported quarantine operations fail per-operation without aborting later work. +- Removed silent duplicate-resolution fallback: unresolved `by_quality` groups and unsupported strategies now surface as explicit manual-review outcomes. +- Canonicalized duplicate path matching so duplicate analysis joins stay stable across path-format differences. +- Defined scanner behavior for non-zero `find` exits: keep partial stdout with a warning, or return an empty deterministic result if no paths were emitted. +- Moved plan loading onto a validated typed `ExecutionPlan` boundary in the I/O layer. +- Updated repository review artifacts and documentation to reflect the new baseline. +- Verification: `pytest -q` → `507 passed`. + ## 2026-04-02 ### Schema Validation & Deterministic Planning diff --git a/CLAUDE.md b/CLAUDE.md index 8f88287..e95e908 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,7 +1,7 @@ # CLAUDE.md ## Documentation Status -- Synchronized with refactor baseline on 2026-02-16 (see `CHANGELOG.md`). +- Synchronized with post-refactor baseline on 2026-04-07 (see `CHANGELOG.md`). This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. @@ -25,6 +25,8 @@ uv pip install -e ".[dev]" uv pip install -e ".[tui]" ``` +The default CLI path does not require Textual; it is imported only when `uv run vlm review-plan --tui` is used. + ### Testing ```bash # Run all tests (use uv run to execute in the project environment) @@ -52,17 +54,20 @@ uv run vlm --help uv run vlm config init # Common workflow -uv run vlm scan # Discover files -uv run vlm parse # Extract identities (v1 schema) -uv run vlm parse --inventory inventory.csv # Extract identities with embedded metadata (v2 schema, recommended) -uv run vlm enrich # (Optional) Enrich titles/reputation via TMDB -uv run vlm enrich --refresh-all # Force full refresh (ignore cache) -uv run vlm analyze # Detect gaps/duplicates -uv run vlm plan # Generate execution plan -uv run vlm plan --analysis analysis.json # Generate plan with duplicate resolution -uv run vlm execute # Dry-run (default) -uv run vlm execute --confirm # Actually execute -uv run vlm rollback # Undo executed operations +uv run vlm scan # Discover files +uv run vlm parse # Extract identities (v1 schema) +uv run vlm parse --inventory artifacts/inventory.csv # Embed metadata for quality-aware duplicate handling +uv run vlm enrich # (Optional) Enrich titles/reputation via TMDB +uv run vlm enrich --refresh-all # Force full refresh (ignore cache) +uv run vlm analyze # Detect gaps/duplicates +uv run vlm plan # Generate execution plan +uv run vlm plan --analysis artifacts/analysis.json # Generate plan with duplicate resolution +uv run vlm review-plan # Export CSV + terminal review preview +uv run vlm review-plan --tui # Optional full-screen review UI +uv run vlm apply-review # Sync edited CSV decisions back into the plan +uv run vlm execute # Dry-run (default) +uv run vlm execute --confirm # Actually execute +uv run vlm rollback # Undo executed operations # Reporting uv run vlm report summary # Overview statistics @@ -84,15 +89,16 @@ uv run vlm state set --status reviewed ### Core Workflow VLM follows a read-first, multi-stage pipeline: -1. **Scan** → discovers video files, extracts metadata via ffprobe (optional), saves to inventory.csv -2. **Parse** → extracts titles/years/seasons/episodes from filenames, saves to identities.json - - Use `--inventory inventory.csv` to embed video metadata (v2 schema) for accurate duplicate resolution by quality +1. **Scan** → discovers video files, extracts metadata via ffprobe (optional), saves to `artifacts/inventory.csv` +2. **Parse** → extracts titles/years/seasons/episodes from filenames, saves to `artifacts/identities.json` + - Use `--inventory artifacts/inventory.csv` to embed video metadata (v2 schema) for accurate duplicate resolution by quality - Without `--inventory`, produces v1 schema (lightweight, no embedded metadata) -3. **Enrich** (optional) → adds bilingual titles and reputation (TMDB); updates identities.json in place; uses SQLite cache for incremental runs -4. **Analyze** → detects episode gaps and duplicates, saves to analysis.json -5. **Plan** → generates reviewable execution plan (plan.json) with file operations -6. **Execute** → performs file operations (dry-run by default, --confirm to execute) -7. **Rollback** → reverses executed operations (best-effort) +3. **Enrich** (optional) → adds bilingual titles and reputation (TMDB); updates `artifacts/identities.json` in place; uses SQLite cache for incremental runs +4. **Analyze** → detects episode gaps and duplicates, saves to `artifacts/analysis.json` +5. **Plan** → generates reviewable execution plan (`artifacts/plan.json`) with file operations +6. **Review / Apply Review** → previews high-risk operations in terminal or optional TUI, then syncs edited CSV decisions back into the plan when needed +7. **Execute** → performs file operations (dry-run by default, `--confirm` to execute) +8. **Rollback** → reverses executed operations (best-effort) ### Module Organization - `cli.py` - Click-based CLI interface, global options, command registration @@ -101,12 +107,12 @@ VLM follows a read-first, multi-stage pipeline: - `scanner.py` - File discovery using system `find` command, metadata extraction via ffprobe - `parser.py` - Filename parsing using regex patterns (movies: title + year, series: SxxExx) - `enrichment.py` - Enrichment pipeline; `cache.py` - SQLite cache; `providers/` - TMDB etc. -- `io.py` - Unified JSON/CSV I/O helpers (including analysis writer and data adapters) +- `io.py` - Unified JSON/CSV I/O helpers, including validated typed plan loading/saving - `utils.py` - UTC time, format_size, shared helpers - `analysis.py` - Completeness checking (episode gaps) and duplicate detection -- `duplicate_resolve.py` - Duplicate group resolution (by_quality, by_reputation, first_seen, manual) -- `planner.py` - Execution plan generation with conflict detection -- `executor.py` - File operations (move/rename/quarantine) with rollback logging +- `duplicate_resolve.py` - Duplicate group resolution with explicit failures for unresolved strategies/data +- `planner.py` - Execution plan generation with conflict detection and manual-review duplicate fallback +- `executor.py` - File operations (move/rename/quarantine) with rollback logging and library-root safety checks - `quarantine.py` - Quarantine management with manifest tracking - `state.py` - File state tracking across workflow stages - `reports.py` - Report generation (inventory, completeness, duplicates, summary) diff --git a/CODE_IMPROVEMENTS.md b/CODE_IMPROVEMENTS.md index 6671ec4..49b3b86 100644 --- a/CODE_IMPROVEMENTS.md +++ b/CODE_IMPROVEMENTS.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # VLM 代码改进清单 diff --git a/FIX_PLAN.md b/FIX_PLAN.md index 24d6f0d..85e3d3b 100644 --- a/FIX_PLAN.md +++ b/FIX_PLAN.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # Fix Plan (Verified Issues Only) diff --git a/GEMINI.md b/GEMINI.md index b793c72..8862a6a 100644 --- a/GEMINI.md +++ b/GEMINI.md @@ -1,7 +1,7 @@ # GEMINI.md ## Documentation Status -- Synced with repository refactor baseline on 2026-02-20 (source of truth: `CHANGELOG.md`). +- Synced with the post-refactor baseline on 2026-04-07 (source of truth: `CHANGELOG.md`). This document provides a comprehensive overview of the Video Library Manager (VLM) project, intended to be used as instructional context for Gemini. @@ -14,8 +14,8 @@ The Video Library Manager (VLM) is a Python-based CLI tool designed for managing * **Scanning & Parsing:** Discovers video files, extracts metadata. Proactively detects `ffprobe` for video properties with graceful fallback to file-level metadata. Parses filenames for titles, years, seasons, and episodes, with specific support for Anime-style hyphenated numbering and release group stripping. * **Metadata Enrichment:** Augments local data with information from TMDB, including bilingual titles and reputation scores. It uses a local SQLite cache to improve performance. * **Analysis:** Detects duplicate files (with quality comparisons) and identifies gaps in TV series episodes. -* **Planning & Execution:** Generates a reviewable JSON-based execution plan for file operations (move, rename, quarantine). Supports "safe mode" and directory preservation. -* **Plan Review Cycle:** Exports high-risk operations to CSV for manual confirmation (`review-plan`) and synchronizes user decisions back to the master plan (`apply-review`), enabling a full human-in-the-loop workflow. +* **Planning & Execution:** Generates a reviewable JSON-based execution plan for file operations (move, rename, quarantine), validates move/rename sources and destinations against `library_root` during execution, and keeps execution failures isolated per operation. +* **Plan Review Cycle:** Exports high-risk operations to CSV for manual confirmation (`review-plan`), offers an optional Textual TUI via `review-plan --tui`, and synchronizes user decisions back to the master plan (`apply-review`). * **Quarantine Management:** Safely isolates files for review, with full support for listing and restoration. * **Reporting:** Creates reports for inventory, duplicate files, and series completeness. * **State Management:** Tracks the status of files throughout the organization workflow. @@ -36,16 +36,15 @@ The project follows a modular structure located in the `src/vlm` directory. * `cli.py`: The main entry point for the CLI, using Click. * `commands/*.py`: Implementation of the individual CLI commands (scan, parse, enrich, analyze, plan, execute/rollback). * `scanner.py`, `parser.py`, `enrichment.py`, `analysis.py`, `planner.py`, `executor.py`: Core logic for the different stages of the workflow. -* `io.py`: Unified I/O layer for JSON and CSV handling. +* `io.py`: Unified I/O layer for JSON and CSV handling, including validated typed plan loading. * `cache.py`: Local SQLite cache for TMDB metadata. * `context.py`: CLI context and state management for command execution. -* `duplicate_resolve.py`: Logic for resolving duplicate files based on quality and metadata. +* `duplicate_resolve.py`: Logic for resolving duplicate files with explicit failure on unsupported or ambiguous inputs. * `logging_config.py`: Centralized logging configuration. * `plan_review.py`: Risk analysis and manual review generation for execution plans. * `quarantine.py`: Management of quarantined files (listing, adding, restoring). * `reports.py`: Generation of inventory, completeness, and duplicate reports. * `state.py`: File status tracking and persistence (reviewed, ignored, planned, etc.). -* `transaction.py`: Atomic filesystem operations and transaction logging for reliability. * `providers/tmdb.py`: Client for interacting with the TMDB API. * `models.py`: Defines the data structures used throughout the application. * `config.py`: Manages application configuration from a YAML file. @@ -59,6 +58,7 @@ The project uses `uv` for dependency management. * Install dependencies: `uv pip install -e .` * Install development dependencies: `uv pip install -e ".[dev]"` +* Install the optional Textual review UI: `uv pip install -e ".[tui]"` **Running the application:** @@ -71,6 +71,8 @@ The main entry point is the `vlm` command. * Analyze the library: `uv run vlm analyze` * Generate a plan: `uv run vlm plan` * Review a plan: `uv run vlm review-plan` +* Review a plan in the optional TUI: `uv run vlm review-plan --tui` +* Apply edited review CSV decisions: `uv run vlm apply-review` * Execute the plan (dry-run): `uv run vlm execute` * Execute the plan (with confirmation): `uv run vlm execute --confirm` * Rollback the last execution: `uv run vlm rollback` diff --git a/IMPLEMENTATION_PLAN_2026-02-13.md b/IMPLEMENTATION_PLAN_2026-02-13.md index de3c42a..1d7d239 100644 --- a/IMPLEMENTATION_PLAN_2026-02-13.md +++ b/IMPLEMENTATION_PLAN_2026-02-13.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # VLM 改进实施计划(2026-02-13) diff --git a/IMPROVEMENT_RECOMMENDATIONS_2026-02-13.md b/IMPROVEMENT_RECOMMENDATIONS_2026-02-13.md index c6a32fa..9beb368 100644 --- a/IMPROVEMENT_RECOMMENDATIONS_2026-02-13.md +++ b/IMPROVEMENT_RECOMMENDATIONS_2026-02-13.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # VLM 项目与 Skill 改进建议(2026-02-13) diff --git a/README.md b/README.md index 2f756e8..bc8fd84 100644 --- a/README.md +++ b/README.md @@ -2,11 +2,14 @@ ## Documentation Status -- Last synchronized: **2026-04-02** +- Last synchronized: **2026-04-07** +- Validation baseline: **`pytest -q` → 507 passed** - Human-in-the-loop workflow includes `vlm apply-review` for syncing manual plan edits from CSV back into `plan.json`. -- Scanner now detects `ffprobe` availability and degrades gracefully. -- JSON artifacts (`identities.json`, `analysis.json`, `plan.json`) are schema-validated on load/save. +- Scanner now detects `ffprobe` availability, degrades gracefully, and treats partial `find` output as a warning-backed partial result. +- JSON artifacts (`identities.json`, `analysis.json`, `plan.json`) are schema-validated on load/save, and plan loading now crosses a typed `ExecutionPlan` boundary. - Plan generation separates logical intent from environment-derived validation snapshots, improving reproducibility. +- `vlm review-plan --tui` remains an optional Textual dependency; the default CLI path does not require Textual. +- Execution rejects move/rename plans whose source or destination escapes `library_root`. A Python-based CLI tool for managing personal video collections with a safety-first, human-in-the-loop approach. @@ -40,6 +43,8 @@ uv pip install -e ".[dev]" uv pip install -e ".[tui]" ``` +The standard CLI remains fully usable without Textual; the dependency is imported only when `vlm review-plan --tui` is requested. + ## Quick Start ### 1. Initialize Configuration @@ -69,6 +74,7 @@ This creates `artifacts/inventory.csv` with all discovered files and their metad **What happens:** - Discovers video files using the system `find` command +- If `find` exits non-zero after emitting paths, VLM keeps the partial results and logs a warning; if no paths are emitted, discovery returns an empty deterministic result - Extracts file metadata (size, modification time) - Categorizes files based on directory structure (movie/series/anime/other) - Extracts video metadata using ffprobe (if available) @@ -157,6 +163,7 @@ This creates `artifacts/plan.json` with: Safety guards in planning: - Titles used in path templates are sanitized (path separators/control chars/`..` stripped) - Any destination outside `library_root` is rejected as `no-op` +- Duplicate resolution with `--analysis` is deterministic: unresolved `by_quality` groups are converted into explicit manual-review `no-op` operations instead of silently keeping the first file Duplicate keep strategy is configurable in `~/.vlm/config.yaml` under `plan.duplicate_keep`: - `by_quality` - Prefer highest quality (resolution > source > codec > file size). Best for automatic duplicate resolution. @@ -165,9 +172,10 @@ Duplicate keep strategy is configurable in `~/.vlm/config.yaml` under `plan.dupl - `first_seen` - Keep the first file in each duplicate group. - `manual` - Do not generate quarantine operations; duplicates are listed in analysis only. -**Review the plan** in one of three ways: +**Review the plan** in one of four ways: - Open `artifacts/plan.json` in your editor - Run `vlm review-plan` to get a terminal preview (summary + high-risk operation preview) +- Run `vlm review-plan --tui` for the optional full-screen Textual review UI - Run `vlm execute` to see the same plan summary in dry-run mode You can still edit `plan.json` directly when needed. @@ -189,7 +197,8 @@ vlm execute --confirm **Important**: This creates a rollback log in `~/.vlm/rollback/` for reverting changes. Execution safeguards: -- Even if a manually edited `artifacts/plan.json` contains an unsafe destination, execution rejects paths outside `library_root` +- Even if a manually edited `artifacts/plan.json` contains an unsafe move/rename source or destination, execution rejects paths outside `library_root` +- Unsupported quarantine operations are reported as failed results without aborting later operations in the same run - Summary counters treat conflict skips separately from real failures (`failed`/`skipped` are mutually exclusive) ### 8. Rollback (If Needed) diff --git a/TMDB_REFACTOR_PLAN.md b/TMDB_REFACTOR_PLAN.md index 870a6ca..b0c199b 100644 --- a/TMDB_REFACTOR_PLAN.md +++ b/TMDB_REFACTOR_PLAN.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # TMDB Enrichment 重构执行计划 diff --git a/VLM_PROJECT_AUDIT_REPORT.md b/VLM_PROJECT_AUDIT_REPORT.md index f42bf95..9564d63 100644 --- a/VLM_PROJECT_AUDIT_REPORT.md +++ b/VLM_PROJECT_AUDIT_REPORT.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # Video Library Manager (VLM) 深度审计报告 diff --git a/codex_review.md b/codex_review.md index c65acc7..04b0ba2 100644 --- a/codex_review.md +++ b/codex_review.md @@ -1,5 +1,5 @@ > [!NOTE] -> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-02-16). +> Status: Historical snapshot. Current refactor results and validated baseline are tracked in `CHANGELOG.md` (updated 2026-04-07). # Codex 架构复核报告