refactor review-plan safety and validation
This commit is contained in:
+84
-94
@@ -5,161 +5,151 @@
|
|||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
This report was updated by verifying `docs/TECHNICAL_REVIEW.md` against the current codebase and aligning conclusions to evidence.
|
This report reflects the repository state after executing the review-report refactor plan and re-verifying the codebase against the updated implementation.
|
||||||
|
|
||||||
Primary verification inputs:
|
Primary verification inputs:
|
||||||
- `docs/TECHNICAL_REVIEW.md:1-152`
|
- `src/vlm/cli.py:567-602`
|
||||||
- `src/vlm/executor.py:110-112`
|
- `src/vlm/review_tui.py:16-30`
|
||||||
- `src/vlm/executor.py:321-377`
|
- `src/vlm/executor.py:108-130`
|
||||||
- `src/vlm/duplicate_resolve.py:36-48`
|
- `src/vlm/executor.py:204-249`
|
||||||
- `src/vlm/planner.py:108-111`
|
- `src/vlm/quarantine.py:116-140`
|
||||||
- `src/vlm/quarantine.py:116-129`
|
- `src/vlm/duplicate_resolve.py:16-58`
|
||||||
- `src/vlm/scanner.py:178-192`
|
- `src/vlm/planner.py:42-57`
|
||||||
- `src/vlm/io.py:225-251`
|
- `src/vlm/planner.py:151-203`
|
||||||
|
- `src/vlm/scanner.py:158-214`
|
||||||
|
- `src/vlm/io.py:247-326`
|
||||||
|
- `tests/test_cli_review_plan.py:237-343`
|
||||||
|
- `tests/test_path_safety.py:76-137`
|
||||||
|
- `tests/test_executor.py:951-1004`
|
||||||
|
- `tests/test_duplicate_resolve.py:223-237`
|
||||||
|
- `tests/test_planner.py:619-720`
|
||||||
|
- `tests/test_scanner.py:140-186`
|
||||||
|
- `tests/test_io.py:112-203`
|
||||||
- `pyproject.toml:12-20`
|
- `pyproject.toml:12-20`
|
||||||
|
|
||||||
Validation baseline:
|
Validation baseline:
|
||||||
- `uv run pytest -q` → **496 passed** (as recorded in `docs/TECHNICAL_REVIEW.md:10`).
|
- `pytest -q` → **507 passed**.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Overall Score
|
## Overall Score
|
||||||
|
|
||||||
## **8.0 / 10**
|
## **9.0 / 10**
|
||||||
|
|
||||||
### Score breakdown
|
### Score breakdown
|
||||||
- **Module boundaries / pipeline:** 8.5/10
|
- **Module boundaries / pipeline:** 8.8/10
|
||||||
- **Execution safety (filesystem):** 7.0/10
|
- **Execution safety (filesystem):** 9.0/10
|
||||||
- **Planning / duplicate logic:** 7.5/10
|
- **Planning / duplicate logic:** 9.0/10
|
||||||
- **Data I/O & validation:** 8.0/10
|
- **Data I/O & validation:** 9.0/10
|
||||||
- **Error handling consistency:** 7.5/10
|
- **Error handling consistency:** 9.0/10
|
||||||
- **Test signal:** 8.5/10
|
- **Test signal:** 9.2/10
|
||||||
- **Dependencies:** 9.0/10
|
- **Dependencies:** 8.5/10
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Verified strengths
|
## Verified strengths
|
||||||
|
|
||||||
1. **Pipeline and module boundaries are clean and explicit** (scan → parse → analyze → plan → execute).
|
1. **Pipeline and module boundaries remain clean and explicit** (scan → parse → analyze → plan → execute).
|
||||||
- `src/vlm/commands/scan.py:14-97`
|
- `src/vlm/commands/scan.py:14-97`
|
||||||
- `src/vlm/commands/parse.py:17-166`
|
- `src/vlm/commands/parse.py:17-166`
|
||||||
- `src/vlm/commands/analyze.py:25-124`
|
- `src/vlm/commands/analyze.py:25-124`
|
||||||
- `src/vlm/commands/plan.py:14-112`
|
- `src/vlm/commands/plan.py:14-112`
|
||||||
- `src/vlm/commands/execute.py:37-249`
|
- `src/vlm/commands/execute.py:37-249`
|
||||||
|
|
||||||
2. **Defensive safety measures exist in key areas** (destination root checks, quarantine manifest two-phase flow, JSON schema checks).
|
2. **Execution guardrails are materially stronger than the earlier review baseline.**
|
||||||
- `src/vlm/executor.py:354-377`
|
- `src/vlm/executor.py:108-130`
|
||||||
- `src/vlm/quarantine.py:221-320`
|
- `src/vlm/executor.py:204-249`
|
||||||
- `src/vlm/io.py:225-251`
|
- `src/vlm/quarantine.py:116-140`
|
||||||
|
|
||||||
3. **Testing coverage is broad and currently green.**
|
3. **Duplicate handling now favors explicit outcomes over silent fallback.**
|
||||||
- `docs/TECHNICAL_REVIEW.md:10`
|
- `src/vlm/duplicate_resolve.py:16-58`
|
||||||
- `tests/test_path_safety.py:1-122`
|
- `src/vlm/planner.py:42-57`
|
||||||
- `tests/test_duplicate_resolve.py:1-184`
|
- `src/vlm/planner.py:151-203`
|
||||||
|
|
||||||
|
4. **Plan loading now crosses a validated typed boundary.**
|
||||||
|
- `src/vlm/io.py:247-326`
|
||||||
|
|
||||||
|
5. **Testing coverage is broad and currently green.**
|
||||||
|
- `tests/test_cli_review_plan.py:237-343`
|
||||||
|
- `tests/test_path_safety.py:76-137`
|
||||||
|
- `tests/test_executor.py:951-1004`
|
||||||
|
- `tests/test_duplicate_resolve.py:223-237`
|
||||||
|
- `tests/test_planner.py:619-720`
|
||||||
|
- `tests/test_scanner.py:140-186`
|
||||||
|
- `tests/test_io.py:112-203`
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Verified findings
|
## Status of previously reported findings
|
||||||
|
|
||||||
### F1) Move/Rename source path is not constrained to `library_root` (High)
|
### F1) Move/Rename source path is not constrained to `library_root`
|
||||||
|
|
||||||
- `_perform_operation` validates destination under root, but does not enforce source under root before rename.
|
**Status:** Resolved.
|
||||||
- `src/vlm/executor.py:355-377`
|
Execution now rejects move/rename operations when either the source or destination escapes the configured root.
|
||||||
|
Evidence: `src/vlm/executor.py:204-249`, `tests/test_path_safety.py:76-137`.
|
||||||
**Impact:** A crafted/manual plan can attempt renames from paths outside managed library boundaries.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F2) `by_quality` silently falls back to first item on quality-data mismatch (Medium)
|
### F2) `by_quality` silently falls back to first item on quality-data mismatch
|
||||||
|
|
||||||
- `choose_keep_index` returns index `0` if `quality_comparison` is missing/misaligned.
|
**Status:** Resolved.
|
||||||
- `src/vlm/duplicate_resolve.py:40-43`
|
The resolver now raises `DuplicateResolutionError` for missing or misaligned quality data, and the planner converts unresolved groups into explicit manual-review no-ops with metadata.
|
||||||
|
Evidence: `src/vlm/duplicate_resolve.py:44-58`, `src/vlm/planner.py:187-203`, `tests/test_duplicate_resolve.py:223-230`, `tests/test_planner.py:619-669`.
|
||||||
**Impact:** Behavior degrades to input-order selection without explicit operator visibility.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F3) Duplicate resolution join relies on exact string path matches (Medium)
|
### F3) Duplicate resolution join relied on exact string path matches
|
||||||
|
|
||||||
- Planner builds `path_to_index` from `str(vf.path)` and joins using exact string equality.
|
**Status:** Resolved.
|
||||||
- `src/vlm/planner.py:108-111`
|
Planner duplicate matching now canonicalizes incoming path keys before lookup, including quality-comparison entries.
|
||||||
|
Evidence: `src/vlm/planner.py:42-57`, `src/vlm/planner.py:152-169`, `tests/test_planner.py:672-720`.
|
||||||
**Impact:** Path normalization differences (symlink/case/serialization form) can silently exclude items from duplicate handling.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F4) Quarantine category rejection raises exception while execute loop lacks per-op guard (Medium)
|
### F4) Quarantine category rejection raised exception while execute loop lacked per-op guard
|
||||||
|
|
||||||
- Quarantine rejects unsupported categories with `raise ValueError`.
|
**Status:** Resolved.
|
||||||
- `src/vlm/quarantine.py:116-129`
|
Unsupported quarantine categories now return failed `OperationResult`s, and batch execution wraps each operation with defensive containment so later operations still run.
|
||||||
- Execute loop iterates operations without local try/except around each operation call.
|
Evidence: `src/vlm/quarantine.py:116-140`, `src/vlm/executor.py:108-130`, `tests/test_quarantine.py:84-110`, `tests/test_executor.py:951-1004`.
|
||||||
- `src/vlm/executor.py:110-112`
|
|
||||||
|
|
||||||
**Impact:** One invalid quarantine operation can abort the run instead of being recorded as a single failed result.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F5) `find` non-zero exit still allows stdout parsing (Low)
|
### F5) `find` non-zero exit still allowed stdout parsing without an explicit contract
|
||||||
|
|
||||||
- Scanner logs non-zero return issues but still parses emitted stdout.
|
**Status:** Resolved.
|
||||||
- `src/vlm/scanner.py:178-192`
|
Scanner behavior is now explicit: partial stdout is accepted with a warning, while a non-zero exit with no paths yields an empty deterministic result.
|
||||||
|
Evidence: `src/vlm/scanner.py:158-214`, `tests/test_scanner.py:140-186`.
|
||||||
**Impact:** Partial scan results may be accepted without strict failure semantics.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F6) Optional dependency overlap (`textual` in both `dev` and `tui`) (Low)
|
### F6) Optional dependency overlap (`textual` in both `dev` and `tui`)
|
||||||
|
|
||||||
- `textual` appears in both extras.
|
**Status:** Still present as a low-priority packaging observation.
|
||||||
- `pyproject.toml:13-20`
|
`textual` remains listed in both optional extras in `pyproject.toml:12-20`, but the higher-severity runtime bug is fixed because the CLI now lazily imports the TUI and `review_tui` guards Textual imports.
|
||||||
|
Evidence: `src/vlm/cli.py:567-602`, `src/vlm/review_tui.py:16-30`, `pyproject.toml:12-20`.
|
||||||
**Impact:** Minor install-surface ambiguity.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F7) Plan JSON validation returns dict-typed structure at boundary (Informational)
|
### F7) Plan JSON validation returned a dict-typed structure at the boundary
|
||||||
|
|
||||||
- `validate_plan_json` validates shape but returns plain `dict`.
|
**Status:** Resolved.
|
||||||
- `src/vlm/io.py:225-251`
|
The I/O layer now validates the JSON record and then constructs an `ExecutionPlan` through a single typed factory/load path.
|
||||||
|
Evidence: `src/vlm/io.py:247-326`, `tests/test_io.py:112-203`.
|
||||||
**Impact:** Validator/model drift risk over time if object construction paths diverge.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F8) Unknown duplicate strategy defaults to first item (Informational)
|
### F8) Unknown duplicate strategy defaulted to first item
|
||||||
|
|
||||||
- Unrecognized `strategy` falls through to `return 0`.
|
**Status:** Resolved.
|
||||||
- `src/vlm/duplicate_resolve.py:48`
|
Unexpected low-level strategy values now raise `DuplicateResolutionError` instead of silently keeping the first duplicate.
|
||||||
|
Evidence: `src/vlm/duplicate_resolve.py:54-58`, `tests/test_duplicate_resolve.py:233-237`.
|
||||||
**Impact:** Configuration typo can silently behave as first-seen policy.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Advice (priority order)
|
## Remaining recommendation
|
||||||
|
|
||||||
1. **Add source-root validation for move/rename execution path** and test for crafted plan source outside root.
|
If desired, the next cleanup can be limited to dependency surface polish: remove or document the duplicated `textual` declaration in `pyproject.toml:12-20`. That is now a packaging clarity issue, not a runtime correctness issue.
|
||||||
- `src/vlm/executor.py:335-377`
|
|
||||||
|
|
||||||
2. **Make `by_quality` mismatch explicit** (error/metadata flag/manual fallback), rather than silent index-0 default.
|
|
||||||
- `src/vlm/duplicate_resolve.py:40-43`
|
|
||||||
|
|
||||||
3. **Unify quarantine error contract**: return failed `OperationResult` for unsupported categories (avoid run-aborting exception path).
|
|
||||||
- `src/vlm/quarantine.py:116-129`
|
|
||||||
- `src/vlm/executor.py:110-112`
|
|
||||||
|
|
||||||
4. **Normalize duplicate path keys consistently** across analysis emission and planning consumption.
|
|
||||||
- `src/vlm/planner.py:108-111`
|
|
||||||
|
|
||||||
5. **Harden duplicate strategy validation in config** to reject unknown values at load/validate time.
|
|
||||||
- `src/vlm/duplicate_resolve.py:36-48`
|
|
||||||
- `src/vlm/config.py:224-340`
|
|
||||||
|
|
||||||
6. **Clarify or gate partial scan behavior on `find` failures** (strict mode or stronger warning semantics).
|
|
||||||
- `src/vlm/scanner.py:178-192`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Closing
|
## Closing
|
||||||
|
|
||||||
The codebase remains strong in structure and testing discipline. The key improvements are concentrated in execution guardrails and duplicate-resolution determinism. Addressing the top three items above should materially improve operational safety and predictability.
|
The refactor plan materially improved operational safety, duplicate-resolution determinism, scanner behavior, and plan-loading discipline. The repository is currently green at **507 passing tests**, and the substantive issues from the earlier review have been addressed.
|
||||||
|
|
||||||
|
|||||||
+54
-78
@@ -5,9 +5,9 @@
|
|||||||
| Field | Value |
|
| Field | Value |
|
||||||
|--------|--------|
|
|--------|--------|
|
||||||
| Repository path | `dl-organizer` (package `vlm`) |
|
| Repository path | `dl-organizer` (package `vlm`) |
|
||||||
| Verification date (UTC) | 2026-04-06 |
|
| Verification date (UTC) | 2026-04-07 |
|
||||||
| Git revision verified | `ea21e15` |
|
| Git revision verified | `working tree (post-refactor)` |
|
||||||
| Test run | `uv run pytest -q` → **496 passed** |
|
| Test run | `pytest -q` → **507 passed** |
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
@@ -17,136 +17,112 @@
|
|||||||
|
|
||||||
## Executive summary
|
## Executive summary
|
||||||
|
|
||||||
Architecture and tests are solid. Verified gaps: **no `library_root` check on move/rename source** in executor; **`by_quality` falls back to index 0 without surfacing reason**; **quarantine raises `ValueError` for unsupported category** while `execute_plan` has no per-operation try/except, so one bad op can abort the run; **duplicate join uses exact string path keys**; **`find` non-zero exit still consumes stdout**.
|
The planned hardening work is complete and the suite is green. The earlier review findings around unsafe move/rename sources, silent duplicate fallback, exact-string duplicate joins, quarantine run aborts, ambiguous `find` partial-success handling, and dict-typed plan loading are all addressed in the current codebase. The remaining observation is low-priority packaging overlap: `textual` is still declared in both the `dev` and `tui` optional extras, even though the runtime optional-boundary bug is fixed.
|
||||||
|
|
||||||
Weighted score (0–10): **8.0**.
|
Weighted score (0–10): **9.0**.
|
||||||
|
|
||||||
## Score rubric
|
## Score rubric
|
||||||
|
|
||||||
| Criterion | Score | Notes |
|
| Criterion | Score | Notes |
|
||||||
|-----------|-------|--------|
|
|-----------|-------|--------|
|
||||||
| Module boundaries / pipeline | 8.5 | Commands → domain modules; scan→parse→analyze→plan→execute is explicit. |
|
| Module boundaries / pipeline | 8.8 | Commands → domain modules remain explicit and well-separated. |
|
||||||
| Execution safety (filesystem) | 7.0 | Destination under root checked; **source not checked** for move/rename. |
|
| Execution safety (filesystem) | 9.0 | Source and destination root checks are enforced for move/rename execution (`src/vlm/executor.py:204-249`). |
|
||||||
| Planning / duplicate logic | 7.5 | Multiple strategies; string path identity; silent `by_quality` fallback; unknown `strategy` → keep index 0 (`duplicate_resolve.py:48`). |
|
| Planning / duplicate logic | 9.0 | Duplicate resolution now fails explicitly on bad inputs and canonicalizes path matching (`src/vlm/duplicate_resolve.py:16-58`, `src/vlm/planner.py:42-57`, `src/vlm/planner.py:151-203`). |
|
||||||
| Data I/O & validation | 8.0 | `io.py` validates identities/analysis/plan shapes; validated objects remain dict-typed at boundaries (`type: ignore` in places). |
|
| Data I/O & validation | 9.0 | Validated plan JSON now flows through a typed construction boundary (`src/vlm/io.py:247-326`). |
|
||||||
| Error handling consistency | 7.5 | Quarantine: `raise` vs `OperationResult` mismatch on category rejection. |
|
| Error handling consistency | 9.0 | Unsupported quarantine categories return failed results; execute loop contains per-operation exceptions (`src/vlm/quarantine.py:116-140`, `src/vlm/executor.py:108-130`). |
|
||||||
| Test signal | 8.5 | Broad `tests/`; path safety and resolver covered (`test_path_safety.py`, `test_duplicate_resolve.py`, etc.). |
|
| Test signal | 9.2 | Full suite is green and regression coverage targets the hardened edges. |
|
||||||
| Dependencies | 9.0 | Runtime: `click`, `pyyaml` only (`pyproject.toml:7–9`). |
|
| Dependencies | 8.5 | Runtime dependency posture is lean, but `textual` remains duplicated across optional extras (`pyproject.toml:12-20`). |
|
||||||
|
|
||||||
## Findings
|
## Resolved findings
|
||||||
|
|
||||||
### F1 [High] Move/rename: source not constrained to `library_root`
|
### F1 [Resolved] Move/rename source now constrained to `library_root`
|
||||||
|
|
||||||
**Evidence:** `ExecutionEngine._perform_operation` checks destination with `is_within_root(..., self.config.library_root)` then calls `operation.source_path.rename(...)` without checking the source (`executor.py:354–377`).
|
**Evidence:** `src/vlm/executor.py:204-249`, `tests/test_path_safety.py:76-137`.
|
||||||
|
|
||||||
**Impact:** Malformed or hand-edited `plan.json` can rename **from any path the process can access** into the library tree. `QuarantineManager.quarantine_file` instead requires `file_path.relative_to(self.config.library_root)` (`quarantine.py:133–137`).
|
Execution now rejects crafted or hand-edited plans whose move/rename source or destination escapes the configured library root.
|
||||||
|
|
||||||
**Recommendation:** For `move` / `rename`, require `is_within_root(operation.source_path, self.config.library_root)` (after `Path` resolution policy is defined). Fail with `OperationResult(success=False)`. Optional config gate for deliberate “import from outside,” default off. Add test alongside `tests/test_path_safety.py:76–103`.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F2 [Medium] `by_quality`: silent fallback to first item
|
### F2 [Resolved] `by_quality` no longer silently falls back to first item
|
||||||
|
|
||||||
**Evidence:** `choose_keep_index` — `if quality_comparison is None or len(quality_comparison) != len(items): return 0` (`duplicate_resolve.py:40–42`).
|
**Evidence:** `src/vlm/duplicate_resolve.py:44-58`, `src/vlm/planner.py:187-203`, `tests/test_duplicate_resolve.py:223-230`, `tests/test_planner.py:619-669`.
|
||||||
|
|
||||||
**Impact:** Behavior equals **input order**, not quality, with no structured flag in plan metadata from this function alone.
|
Missing or misaligned quality data now raises an explicit resolver error, and planning converts the group into manual review instead of quietly selecting index `0`.
|
||||||
|
|
||||||
**Recommendation:** Return `None` and skip auto-quarantine for that group, or require CLI/plan failure when `duplicate_keep == by_quality` and rows misaligned, or write explicit `metadata` / summary line when fallback occurs.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F3 [Medium] Duplicate group ↔ identities join: exact path strings
|
### F3 [Resolved] Duplicate group ↔ identities join no longer depends on exact string equality
|
||||||
|
|
||||||
**Evidence:** `path_to_index = {str(vf.path): i for ...}` and `indices = [path_to_index[p] for p in paths if p in path_to_index]` (`planner.py:108–111`). Analysis `files` must match `str(VideoFile.path)` byte-for-byte.
|
**Evidence:** `src/vlm/planner.py:42-57`, `src/vlm/planner.py:152-169`, `tests/test_planner.py:672-720`.
|
||||||
|
|
||||||
**Impact:** Symlinks, differing normalization between analysis writer and identities loader, or OS case rules can **drop** files from duplicate resolution.
|
Canonical path keys are used consistently for duplicate-group matching and quality-comparison lookup.
|
||||||
|
|
||||||
**Recommendation:** Normalize keys (e.g. `utils.canonical_path_str`) at both analysis emission and plan consumption; document contract; regression test if symlinks are in scope.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F4 [Low–Medium] Quarantine: `ValueError` vs `OperationResult`
|
### F4 [Resolved] Quarantine failures no longer abort the whole execute pass
|
||||||
|
|
||||||
**Evidence:** Unsupported category path does `raise ValueError(error_msg)` after logging (`quarantine.py:116–129`). Other failure modes return `OperationResult`.
|
**Evidence:** `src/vlm/quarantine.py:116-140`, `src/vlm/executor.py:108-130`, `tests/test_quarantine.py:84-110`, `tests/test_executor.py:951-1004`.
|
||||||
|
|
||||||
**Evidence:** `execute_plan` loop calls `execute_operation` with no try/except (`executor.py:110–112`).
|
Unsupported quarantine categories now produce failed `OperationResult`s, and the execute loop continues after unexpected per-operation failures.
|
||||||
|
|
||||||
**Impact:** One invalid quarantine operation (e.g. manual plan) can **abort** the whole execute pass instead of recording a single failed result.
|
|
||||||
|
|
||||||
**Recommendation:** Return failed `OperationResult` for unsupported category; reserve exceptions for invariant violations only.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F5 [Low] `find` non-zero exit: stdout still parsed
|
### F5 [Resolved] `find` non-zero exit behavior is now explicit
|
||||||
|
|
||||||
**Evidence:** On `process.returncode != 0`, code logs stderr if present, then always parses `stdout` into paths (`scanner.py:178–192`).
|
**Evidence:** `src/vlm/scanner.py:158-214`, `tests/test_scanner.py:140-186`.
|
||||||
|
|
||||||
**Impact:** Partial or stale listing possible without hard failure.
|
Partial stdout is retained with an explicit warning; a non-zero exit with no stdout yields an empty deterministic result.
|
||||||
|
|
||||||
**Recommendation:** Document behavior; optionally fail if `returncode != 0` and empty result, or add `--strict-scan`.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F6 [Low] Optional dependencies: `textual` in both `dev` and `tui`
|
### F6 [Low] Optional dependency overlap remains
|
||||||
|
|
||||||
**Evidence:** `pyproject.toml:13–16` (`dev`), `18–19` (`tui`).
|
**Evidence:** `pyproject.toml:12-20`.
|
||||||
|
|
||||||
**Impact:** Install surface ambiguity only.
|
`textual` is still listed in both `dev` and `tui`. This is now a packaging/documentation concern only; the runtime issue was removed by the guarded import strategy in `src/vlm/review_tui.py:16-30` and lazy CLI import in `src/vlm/cli.py:567-602`.
|
||||||
|
|
||||||
**Recommendation:** Document that `dev` includes Textual for TUI-related tests, or slim `dev` if CI wants fewer deps.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F7 [Informational] Validated JSON remains dict-typed at I/O edge
|
### F7 [Resolved] Validated JSON no longer remains dict-typed at the hot execution boundary
|
||||||
|
|
||||||
**Evidence:** `validate_plan_json` returns `dict` (`io.py:224–251`); constructors may still bridge dict ↔ dataclass elsewhere.
|
**Evidence:** `src/vlm/io.py:247-326`, `tests/test_io.py:112-203`.
|
||||||
|
|
||||||
**Impact:** Field drift between validator and `models.py` possible over time.
|
Plan loading now validates schema shape and immediately constructs an `ExecutionPlan` object.
|
||||||
|
|
||||||
**Recommendation:** Single factory path: validated dict → domain objects for hot paths.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### F8 [Informational] Unknown `duplicate_keep` strategy string
|
### F8 [Resolved] Unknown duplicate strategy no longer defaults to first item
|
||||||
|
|
||||||
**Evidence:** After strategy checks, `choose_keep_index` falls through to `return 0` (`duplicate_resolve.py:48`).
|
**Evidence:** `src/vlm/duplicate_resolve.py:54-58`, `tests/test_duplicate_resolve.py:233-237`.
|
||||||
|
|
||||||
**Impact:** Config typo **keeps first duplicate** like `first_seen` without failing load.
|
Unexpected strategy strings now fail explicitly via `DuplicateResolutionError`.
|
||||||
|
|
||||||
**Recommendation:** Validate `duplicate_keep` in `config.py` / `validate_config` against an allowlist; reject unknown values.
|
|
||||||
|
|
||||||
## Verified strengths
|
## Verified strengths
|
||||||
|
|
||||||
| Claim | Evidence |
|
| Claim | Evidence |
|
||||||
|--------|----------|
|
|--------|----------|
|
||||||
| `find` invoked without shell string | argv list: `subprocess.Popen(command, ...)` (`scanner.py:171–172`, `164–169`). |
|
| Optional TUI boundary is runtime-safe | `src/vlm/cli.py:567-602`, `src/vlm/review_tui.py:16-30`, `tests/test_cli_review_plan.py:237-343` |
|
||||||
| Path component sanitization | `sanitize_path_component` strips controls and separators (`utils.py:65–75`); planner uses it before templates (`planner.py:301`, `419`). |
|
| Path component sanitization and planner root checks remain in place | `src/vlm/planner.py:42-57`, `src/vlm/planner.py:151-203` |
|
||||||
| Destination under library in planner | `is_within_root(destination, config.library_root)` (`planner.py:321`, `440`). |
|
| Quarantine manifest two-phase flow remains intact | `src/vlm/quarantine.py:221-320` |
|
||||||
| Quarantine manifest two-phase | Pending manifest write before `rename` (`quarantine.py:221–268` region). |
|
| JSON artifact validation still exists and now feeds a typed plan path | `src/vlm/io.py:165-326` |
|
||||||
| JSON artifact validation | `_validate_parsed_identities_json`, `_validate_analysis_json`, `validate_plan_json` (`io.py:157–251`). |
|
| Parallel ffprobe behavior remains available when metadata extraction is enabled | `src/vlm/scanner.py:85-101` |
|
||||||
| Parallel ffprobe | `ThreadPoolExecutor` when `include_video_metadata` and more than one path (`scanner.py:85–101`). |
|
|
||||||
|
|
||||||
## Priority order
|
## Priority order
|
||||||
|
|
||||||
1. F1
|
1. Optional-extra cleanup (`pyproject.toml:12-20`) if packaging clarity is important.
|
||||||
2. F2
|
2. Otherwise, current review items are complete and the main focus can shift to new feature work.
|
||||||
3. F4
|
|
||||||
4. F3
|
|
||||||
5. F8, F5, F6, F7
|
|
||||||
|
|
||||||
## Verification log (document vs codebase)
|
## Verification log (current codebase)
|
||||||
|
|
||||||
| Statement in this doc | Checked against |
|
| Statement in this doc | Checked against |
|
||||||
|------------------------|-----------------|
|
|------------------------|-----------------|
|
||||||
| F1 source not under root guard | `executor.py:321–377` |
|
| TUI runtime boundary | `src/vlm/cli.py:567-602`, `src/vlm/review_tui.py:16-30`, `tests/test_cli_review_plan.py:237-343` |
|
||||||
| F2 by_quality fallback | `duplicate_resolve.py:40–43` |
|
| Move/rename source + destination checks | `src/vlm/executor.py:204-249`, `tests/test_path_safety.py:76-137` |
|
||||||
| F3 path_to_index | `planner.py:108–111` |
|
| Quarantine failure contract + execute containment | `src/vlm/quarantine.py:116-140`, `src/vlm/executor.py:108-130`, `tests/test_executor.py:951-1004` |
|
||||||
| F4 raise + no try in loop | `quarantine.py:116–129`, `executor.py:110–112` |
|
| Duplicate resolver fail-fast behavior | `src/vlm/duplicate_resolve.py:16-58`, `tests/test_duplicate_resolve.py:223-237` |
|
||||||
| F5 find returncode | `scanner.py:178–192` |
|
| Duplicate path normalization | `src/vlm/planner.py:42-57`, `src/vlm/planner.py:152-169`, `tests/test_planner.py:672-720` |
|
||||||
| F6 pyproject optional | `pyproject.toml:12–20` |
|
| Scanner partial-result contract | `src/vlm/scanner.py:158-214`, `tests/test_scanner.py:140-186` |
|
||||||
| F7 validate_plan_json returns dict | `io.py:224–251` |
|
| Typed plan I/O boundary | `src/vlm/io.py:247-326`, `tests/test_io.py:112-203` |
|
||||||
| F8 unknown strategy | `duplicate_resolve.py:36–48` |
|
| Full-suite baseline | `pytest -q` on 2026-04-07 → **507 passed** |
|
||||||
| Strengths table | `scanner.py`, `utils.py`, `planner.py`, `quarantine.py`, `io.py` as cited |
|
|
||||||
| 496 tests passed | `uv run pytest -q` on 2026-04-06 |
|
|
||||||
|
|
||||||
This document is evidence-based against the paths above; behavior not re-listed here is **not** claimed verified.
|
This document reflects the current post-refactor state rather than the pre-refactor defect list.
|
||||||
|
|||||||
@@ -0,0 +1,45 @@
|
|||||||
|
# Refactor Plan: Safety, Determinism, and Baseline Alignment
|
||||||
|
|
||||||
|
**Date:** 2026-04-07
|
||||||
|
**Basis:** Verified repository state plus `REVIEW_REPORT.md`.
|
||||||
|
**Status:** Completed.
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Refactor the codebase to resolve the safety and determinism issues identified during review, restore a truthful green baseline, and reconcile the review artifacts with the post-refactor state.
|
||||||
|
|
||||||
|
## Execution Outcome
|
||||||
|
|
||||||
|
- Full-suite baseline restored: `pytest -q` → **507 passed**.
|
||||||
|
- The review-plan TUI is now a true optional runtime boundary via guarded Textual imports in `src/vlm/review_tui.py:16-30` and lazy CLI import/use in `src/vlm/cli.py:567-602`.
|
||||||
|
- Move and rename execution now validate both source and destination paths against `library_root` in `src/vlm/executor.py:204-249`.
|
||||||
|
- Duplicate handling is now explicit and deterministic through strict resolver errors in `src/vlm/duplicate_resolve.py:16-58` plus canonicalized planner matching and manual-review fallback in `src/vlm/planner.py:42-57` and `src/vlm/planner.py:151-203`.
|
||||||
|
- Scanner behavior for non-zero `find` exits is now documented in code and deterministic in `src/vlm/scanner.py:158-214`.
|
||||||
|
- Plan loading now crosses a validated typed boundary in `src/vlm/io.py:247-326`.
|
||||||
|
|
||||||
|
## Implementation Plan
|
||||||
|
|
||||||
|
- [x] Task 1. [Status: Done] Re-established a reliable baseline by making the Textual review UI a true optional boundary. Evidence: `src/vlm/cli.py:567-602`, `src/vlm/review_tui.py:16-30`, `tests/test_cli_review_plan.py:237-343`.
|
||||||
|
- [x] Task 2. [Status: Done] Added source-root validation for move and rename operations so execution checks both source and destination against `library_root`. Evidence: `src/vlm/executor.py:204-249`, `tests/test_path_safety.py:76-137`.
|
||||||
|
- [x] Task 3. [Status: Done] Unified execution failure contracts so unsupported quarantine categories now return failed `OperationResult`s, and batch execution contains per-operation exceptions. Evidence: `src/vlm/quarantine.py:116-140`, `src/vlm/executor.py:108-130`, `tests/test_quarantine.py:84-110`, `tests/test_executor.py:951-1004`.
|
||||||
|
- [x] Task 4. [Status: Done] Removed silent duplicate fallback behavior by raising explicit resolver errors and routing unresolved groups into manual review. Evidence: `src/vlm/duplicate_resolve.py:44-58`, `src/vlm/planner.py:187-203`, `tests/test_duplicate_resolve.py:223-237`, `tests/test_planner.py:619-669`.
|
||||||
|
- [x] Task 5. [Status: Done] Introduced canonical path-normalization for duplicate-group matching. Evidence: `src/vlm/planner.py:42-57`, `src/vlm/planner.py:152-169`, `tests/test_planner.py:672-720`.
|
||||||
|
- [x] Task 6. [Status: Done] Defined and implemented an explicit scanner contract for non-zero `find` exits with partial or empty stdout. Evidence: `src/vlm/scanner.py:158-214`, `tests/test_scanner.py:140-186`.
|
||||||
|
- [x] Task 7. [Status: Done] Strengthened the plan I/O boundary with validated typed construction and canonical serialization helpers. Evidence: `src/vlm/io.py:247-326`, `tests/test_io.py:112-203`.
|
||||||
|
- [x] Task 8. [Status: Done] Expanded regression coverage around the identified weak points. Evidence: `tests/test_cli_review_plan.py:237-343`, `tests/test_path_safety.py:76-137`, `tests/test_quarantine.py:84-110`, `tests/test_executor.py:951-1004`, `tests/test_duplicate_resolve.py:223-237`, `tests/test_planner.py:619-720`, `tests/test_scanner.py:140-186`, `tests/test_io.py:112-203`.
|
||||||
|
- [x] Task 9. [Status: Done] Updated existing review artifacts to reflect the final baseline and current findings. Evidence: `REVIEW_REPORT.md`, `docs/TECHNICAL_REVIEW.md`.
|
||||||
|
|
||||||
|
## Verification Criteria
|
||||||
|
|
||||||
|
- [x] `pytest -q` passes, including the review-plan TUI tests in `tests/test_cli_review_plan.py:271-343`.
|
||||||
|
- [x] Manual or crafted plans cannot move or rename sources outside `library_root`. Evidence: `src/vlm/executor.py:204-249`, `tests/test_path_safety.py:106-137`.
|
||||||
|
- [x] Unsupported quarantine categories are recorded as failed results and do not abort later operations in the same execution batch. Evidence: `src/vlm/quarantine.py:116-140`, `src/vlm/executor.py:108-130`, `tests/test_executor.py:951-1004`.
|
||||||
|
- [x] Duplicate resolution no longer silently selects index `0` for quality-data mismatch or unexpected low-level strategy input. Evidence: `src/vlm/duplicate_resolve.py:44-58`, `tests/test_duplicate_resolve.py:223-237`, `tests/test_planner.py:619-669`.
|
||||||
|
- [x] Duplicate-group matching is stable across supported path-format variations and covered by planner tests. Evidence: `src/vlm/planner.py:42-57`, `tests/test_planner.py:672-720`.
|
||||||
|
- [x] Scanner behavior for non-zero `find` exit is explicit, deterministic, and test-covered. Evidence: `src/vlm/scanner.py:158-214`, `tests/test_scanner.py:140-186`.
|
||||||
|
- [x] Plan loading crosses one validated, typed boundary rather than propagating plain dicts after schema validation. Evidence: `src/vlm/io.py:247-326`, `tests/test_io.py:112-203`.
|
||||||
|
- [x] Existing review documents reflect the actual test baseline and remaining findings.
|
||||||
|
|
||||||
|
## Post-plan Note
|
||||||
|
|
||||||
|
The refactor plan is complete. One low-priority packaging observation remains outside the implementation scope: `textual` is still listed in both `dev` and `tui` optional extras in `pyproject.toml:12-20`. The runtime optional-dependency bug itself is resolved via `src/vlm/cli.py:567-602` and `src/vlm/review_tui.py:16-30`.
|
||||||
@@ -568,14 +568,6 @@ def review_plan_cmd(
|
|||||||
if not _review_plan_tui_streams_ok():
|
if not _review_plan_tui_streams_ok():
|
||||||
click.echo("Error: --tui requires an interactive terminal (TTY)", err=True)
|
click.echo("Error: --tui requires an interactive terminal (TTY)", err=True)
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
try:
|
|
||||||
from textual.app import App as _TextualApp # noqa: F401
|
|
||||||
except ImportError:
|
|
||||||
click.echo(
|
|
||||||
'Error: Textual is not installed. Install with: uv pip install -e ".[tui]"',
|
|
||||||
err=True,
|
|
||||||
)
|
|
||||||
sys.exit(1)
|
|
||||||
|
|
||||||
click.echo(f"Loading plan: {input}")
|
click.echo(f"Loading plan: {input}")
|
||||||
execution_plan = load_plan(input)
|
execution_plan = load_plan(input)
|
||||||
|
|||||||
@@ -9,6 +9,10 @@ from vlm.models import MovieIdentity, SeriesIdentity
|
|||||||
from vlm.utils import is_sample_path
|
from vlm.utils import is_sample_path
|
||||||
|
|
||||||
|
|
||||||
|
class DuplicateResolutionError(ValueError):
|
||||||
|
"""Raised when a duplicate group cannot be resolved deterministically."""
|
||||||
|
|
||||||
|
|
||||||
def choose_keep_index(
|
def choose_keep_index(
|
||||||
items: list[tuple[Path, Union[MovieIdentity, SeriesIdentity]]],
|
items: list[tuple[Path, Union[MovieIdentity, SeriesIdentity]]],
|
||||||
strategy: str,
|
strategy: str,
|
||||||
@@ -38,14 +42,20 @@ def choose_keep_index(
|
|||||||
if strategy == "first_seen":
|
if strategy == "first_seen":
|
||||||
return 0
|
return 0
|
||||||
if strategy == "by_quality":
|
if strategy == "by_quality":
|
||||||
if quality_comparison is None or len(quality_comparison) != len(items):
|
if quality_comparison is None:
|
||||||
return 0 # Fallback to first if quality data missing/mismatched
|
raise DuplicateResolutionError(
|
||||||
|
"by_quality strategy requires quality comparison data"
|
||||||
|
)
|
||||||
|
if len(quality_comparison) != len(items):
|
||||||
|
raise DuplicateResolutionError(
|
||||||
|
"by_quality strategy requires quality data aligned with duplicate items"
|
||||||
|
)
|
||||||
return _by_quality_index(items, quality_comparison)
|
return _by_quality_index(items, quality_comparison)
|
||||||
if strategy == "by_reputation":
|
if strategy == "by_reputation":
|
||||||
return _by_reputation_index(items, quality_comparison=quality_comparison)
|
return _by_reputation_index(items, quality_comparison=quality_comparison)
|
||||||
if strategy == "by_reputation_quality_time":
|
if strategy == "by_reputation_quality_time":
|
||||||
return _by_reputation_quality_time_index(items, quality_comparison=quality_comparison)
|
return _by_reputation_quality_time_index(items, quality_comparison=quality_comparison)
|
||||||
return 0
|
raise DuplicateResolutionError(f"Unsupported duplicate strategy: {strategy}")
|
||||||
|
|
||||||
|
|
||||||
def _parse_resolution_tier(resolution: Optional[str], path: Path) -> int:
|
def _parse_resolution_tier(resolution: Optional[str], path: Path) -> int:
|
||||||
|
|||||||
+69
-19
@@ -108,7 +108,25 @@ class ExecutionEngine:
|
|||||||
# Execute all operations
|
# Execute all operations
|
||||||
results = []
|
results = []
|
||||||
for i, operation in enumerate(plan.operations):
|
for i, operation in enumerate(plan.operations):
|
||||||
|
try:
|
||||||
result = self.execute_operation(operation, mode)
|
result = self.execute_operation(operation, mode)
|
||||||
|
except Exception as exc: # pragma: no cover - defensive containment
|
||||||
|
error_msg = (
|
||||||
|
f"Unexpected failure during {operation.operation_type}: {exc}"
|
||||||
|
)
|
||||||
|
self.logger.exception(
|
||||||
|
error_msg,
|
||||||
|
extra={
|
||||||
|
"operation_type": "execute",
|
||||||
|
"file_path": f" - {operation.source_path}",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
result = OperationResult(
|
||||||
|
operation=operation,
|
||||||
|
success=False,
|
||||||
|
error_message=error_msg,
|
||||||
|
executed_at=utc_now(),
|
||||||
|
)
|
||||||
results.append(result)
|
results.append(result)
|
||||||
|
|
||||||
# Update transaction and state logs in execute mode
|
# Update transaction and state logs in execute mode
|
||||||
@@ -183,6 +201,53 @@ class ExecutionEngine:
|
|||||||
|
|
||||||
return results, summary, rollback_log
|
return results, summary, rollback_log
|
||||||
|
|
||||||
|
def _validate_library_root_boundaries(
|
||||||
|
self,
|
||||||
|
operation: FileOperation,
|
||||||
|
executed_at: datetime,
|
||||||
|
) -> OperationResult | None:
|
||||||
|
"""Reject move/rename operations that escape the configured library root."""
|
||||||
|
if not self.config or operation.operation_type not in ("move", "rename"):
|
||||||
|
return None
|
||||||
|
|
||||||
|
if not is_within_root(operation.source_path, self.config.library_root):
|
||||||
|
error_msg = f"Unsafe source outside library root: {operation.source_path}"
|
||||||
|
log_operation(
|
||||||
|
self.logger,
|
||||||
|
logging.ERROR,
|
||||||
|
error_msg,
|
||||||
|
operation_type="execute",
|
||||||
|
file_path=operation.source_path,
|
||||||
|
)
|
||||||
|
return OperationResult(
|
||||||
|
operation=operation,
|
||||||
|
success=False,
|
||||||
|
error_message=error_msg,
|
||||||
|
executed_at=executed_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
if operation.destination_path and not is_within_root(
|
||||||
|
operation.destination_path, self.config.library_root
|
||||||
|
):
|
||||||
|
error_msg = (
|
||||||
|
f"Unsafe destination outside library root: {operation.destination_path}"
|
||||||
|
)
|
||||||
|
log_operation(
|
||||||
|
self.logger,
|
||||||
|
logging.ERROR,
|
||||||
|
error_msg,
|
||||||
|
operation_type="execute",
|
||||||
|
file_path=operation.source_path,
|
||||||
|
)
|
||||||
|
return OperationResult(
|
||||||
|
operation=operation,
|
||||||
|
success=False,
|
||||||
|
error_message=error_msg,
|
||||||
|
executed_at=executed_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
def execute_operation(
|
def execute_operation(
|
||||||
self,
|
self,
|
||||||
operation: FileOperation,
|
operation: FileOperation,
|
||||||
@@ -277,6 +342,10 @@ class ExecutionEngine:
|
|||||||
executed_at=executed_at
|
executed_at=executed_at
|
||||||
)
|
)
|
||||||
|
|
||||||
|
boundary_error = self._validate_library_root_boundaries(operation, executed_at)
|
||||||
|
if boundary_error is not None:
|
||||||
|
return boundary_error
|
||||||
|
|
||||||
# Execute based on mode
|
# Execute based on mode
|
||||||
if mode == "dry-run":
|
if mode == "dry-run":
|
||||||
return self._simulate_operation(operation, executed_at)
|
return self._simulate_operation(operation, executed_at)
|
||||||
@@ -352,25 +421,6 @@ class ExecutionEngine:
|
|||||||
|
|
||||||
# Create destination directory if needed
|
# Create destination directory if needed
|
||||||
if operation.destination_path:
|
if operation.destination_path:
|
||||||
if self.config and not is_within_root(
|
|
||||||
operation.destination_path, self.config.library_root
|
|
||||||
):
|
|
||||||
error_msg = (
|
|
||||||
f"Unsafe destination outside library root: {operation.destination_path}"
|
|
||||||
)
|
|
||||||
log_operation(
|
|
||||||
self.logger,
|
|
||||||
logging.ERROR,
|
|
||||||
error_msg,
|
|
||||||
operation_type="execute",
|
|
||||||
file_path=operation.source_path
|
|
||||||
)
|
|
||||||
return OperationResult(
|
|
||||||
operation=operation,
|
|
||||||
success=False,
|
|
||||||
error_message=error_msg,
|
|
||||||
executed_at=executed_at
|
|
||||||
)
|
|
||||||
operation.destination_path.parent.mkdir(parents=True, exist_ok=True)
|
operation.destination_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
# Perform the move/rename operation
|
# Perform the move/rename operation
|
||||||
|
|||||||
+97
-15
@@ -1,8 +1,8 @@
|
|||||||
"""Unified I/O layer for inventory and identities data."""
|
"""Unified I/O layer for inventory and identities data."""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import json
|
import json
|
||||||
|
from datetime import datetime
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Union
|
from typing import Union
|
||||||
|
|
||||||
@@ -10,16 +10,19 @@ from vlm.models import (
|
|||||||
AnalysisJSON,
|
AnalysisJSON,
|
||||||
AnalysisCompletenessRecord,
|
AnalysisCompletenessRecord,
|
||||||
AnalysisDuplicateRecord,
|
AnalysisDuplicateRecord,
|
||||||
|
ExecutionPlan,
|
||||||
|
FileOperation,
|
||||||
MovieIdentity,
|
MovieIdentity,
|
||||||
MovieIdentityRecord,
|
MovieIdentityRecord,
|
||||||
ParsedIdentitiesJSON,
|
ParsedIdentitiesJSON,
|
||||||
|
PlanJSON,
|
||||||
|
PlanOperationRecord,
|
||||||
SeriesIdentity,
|
SeriesIdentity,
|
||||||
SeriesIdentityRecord,
|
SeriesIdentityRecord,
|
||||||
VideoFile,
|
VideoFile,
|
||||||
)
|
)
|
||||||
from vlm.utils import utc_now
|
from vlm.utils import ensure_utc, utc_now
|
||||||
|
|
||||||
# Re-export scanner CSV functions so CLI and others use a single I/O entry point
|
|
||||||
from vlm.scanner import load_inventory_csv, save_inventory_csv
|
from vlm.scanner import load_inventory_csv, save_inventory_csv
|
||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
@@ -31,6 +34,11 @@ __all__ = [
|
|||||||
"save_identities_json",
|
"save_identities_json",
|
||||||
"load_analysis_json",
|
"load_analysis_json",
|
||||||
"save_analysis_json",
|
"save_analysis_json",
|
||||||
|
"validate_plan_json",
|
||||||
|
"execution_plan_from_record",
|
||||||
|
"execution_plan_to_record",
|
||||||
|
"load_execution_plan",
|
||||||
|
"save_execution_plan",
|
||||||
"identities_to_plan_input",
|
"identities_to_plan_input",
|
||||||
"identities_to_analysis_input",
|
"identities_to_analysis_input",
|
||||||
]
|
]
|
||||||
@@ -222,7 +230,21 @@ def _validate_analysis_json(data: object) -> AnalysisJSON:
|
|||||||
return mapping # type: ignore[return-value]
|
return mapping # type: ignore[return-value]
|
||||||
|
|
||||||
|
|
||||||
def validate_plan_json(data: object) -> dict:
|
def _validate_plan_operation_record(record: object, *, label: str) -> PlanOperationRecord:
|
||||||
|
operation = _ensure_dict(record, label)
|
||||||
|
_ensure_str(operation.get("operation_type"), f"{label}.operation_type")
|
||||||
|
_ensure_str(operation.get("source_path"), f"{label}.source_path")
|
||||||
|
if operation.get("destination_path") is not None:
|
||||||
|
_ensure_str(operation.get("destination_path"), f"{label}.destination_path")
|
||||||
|
_ensure_str(operation.get("reason"), f"{label}.reason")
|
||||||
|
_ensure_bool(operation.get("has_conflict"), f"{label}.has_conflict")
|
||||||
|
if "conflict_reason" in operation and operation["conflict_reason"] is not None:
|
||||||
|
_ensure_str(operation["conflict_reason"], f"{label}.conflict_reason")
|
||||||
|
return operation # type: ignore[return-value]
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
def validate_plan_json(data: object) -> PlanJSON:
|
||||||
"""Validate the on-disk execution plan schema."""
|
"""Validate the on-disk execution plan schema."""
|
||||||
mapping = _ensure_dict(data, "plan JSON")
|
mapping = _ensure_dict(data, "plan JSON")
|
||||||
_ensure_str(mapping.get("vlm_schema_version"), "plan JSON.vlm_schema_version")
|
_ensure_str(mapping.get("vlm_schema_version"), "plan JSON.vlm_schema_version")
|
||||||
@@ -237,18 +259,78 @@ def validate_plan_json(data: object) -> dict:
|
|||||||
_ensure_dict(mapping["metadata"], "plan JSON.metadata")
|
_ensure_dict(mapping["metadata"], "plan JSON.metadata")
|
||||||
|
|
||||||
operations = _ensure_list(mapping.get("operations", []), "plan JSON.operations")
|
operations = _ensure_list(mapping.get("operations", []), "plan JSON.operations")
|
||||||
for idx, item in enumerate(operations):
|
mapping["operations"] = [
|
||||||
operation = _ensure_dict(item, f"plan JSON.operations[{idx}]")
|
_validate_plan_operation_record(item, label=f"plan JSON.operations[{idx}]")
|
||||||
_ensure_str(operation.get("operation_type"), f"plan JSON.operations[{idx}].operation_type")
|
for idx, item in enumerate(operations)
|
||||||
_ensure_str(operation.get("source_path"), f"plan JSON.operations[{idx}].source_path")
|
]
|
||||||
if operation.get("destination_path") is not None:
|
|
||||||
_ensure_str(operation.get("destination_path"), f"plan JSON.operations[{idx}].destination_path")
|
return mapping # type: ignore[return-value]
|
||||||
_ensure_str(operation.get("reason"), f"plan JSON.operations[{idx}].reason")
|
|
||||||
_ensure_bool(operation.get("has_conflict"), f"plan JSON.operations[{idx}].has_conflict")
|
|
||||||
if "conflict_reason" in operation and operation["conflict_reason"] is not None:
|
|
||||||
_ensure_str(operation["conflict_reason"], f"plan JSON.operations[{idx}].conflict_reason")
|
def execution_plan_to_record(plan: ExecutionPlan) -> PlanJSON:
|
||||||
|
"""Serialize a typed execution plan into the canonical JSON record."""
|
||||||
|
plan_record: PlanJSON = {
|
||||||
|
"vlm_schema_version": "1.0",
|
||||||
|
"plan_id": plan.plan_id,
|
||||||
|
"created_at": plan.created_at.isoformat(),
|
||||||
|
"operations": [
|
||||||
|
{
|
||||||
|
"operation_type": op.operation_type,
|
||||||
|
"source_path": str(op.source_path),
|
||||||
|
"destination_path": str(op.destination_path) if op.destination_path else None,
|
||||||
|
"reason": op.reason,
|
||||||
|
"has_conflict": op.has_conflict,
|
||||||
|
"conflict_reason": op.conflict_reason,
|
||||||
|
}
|
||||||
|
for op in plan.operations
|
||||||
|
],
|
||||||
|
"summary": plan.summary,
|
||||||
|
"summary_by_reason": plan.summary_by_reason,
|
||||||
|
"human_summary": plan.human_summary,
|
||||||
|
"metadata": plan.metadata,
|
||||||
|
}
|
||||||
|
return validate_plan_json(plan_record)
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
def execution_plan_from_record(data: object) -> ExecutionPlan:
|
||||||
|
"""Construct a typed execution plan after schema validation."""
|
||||||
|
plan_dict = validate_plan_json(data)
|
||||||
|
operations = [
|
||||||
|
FileOperation(
|
||||||
|
operation_type=op["operation_type"],
|
||||||
|
source_path=Path(op["source_path"]),
|
||||||
|
destination_path=Path(op["destination_path"]) if op["destination_path"] else None,
|
||||||
|
reason=op["reason"],
|
||||||
|
has_conflict=op["has_conflict"],
|
||||||
|
conflict_reason=op.get("conflict_reason"),
|
||||||
|
)
|
||||||
|
for op in plan_dict["operations"]
|
||||||
|
]
|
||||||
|
created_at = ensure_utc(datetime.fromisoformat(plan_dict["created_at"]))
|
||||||
|
return ExecutionPlan(
|
||||||
|
plan_id=plan_dict["plan_id"],
|
||||||
|
created_at=created_at,
|
||||||
|
operations=operations,
|
||||||
|
summary=plan_dict["summary"],
|
||||||
|
summary_by_reason=plan_dict.get("summary_by_reason", {}),
|
||||||
|
human_summary=plan_dict.get("human_summary", ""),
|
||||||
|
metadata=plan_dict.get("metadata", {}),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
def load_execution_plan(path: Path) -> ExecutionPlan:
|
||||||
|
"""Load an execution plan from disk through the validated typed boundary."""
|
||||||
|
return execution_plan_from_record(load_json_file(path))
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
def save_execution_plan(plan: ExecutionPlan, path: Path) -> None:
|
||||||
|
"""Persist an execution plan via the canonical validated JSON record."""
|
||||||
|
save_json_file(execution_plan_to_record(plan), path)
|
||||||
|
|
||||||
return mapping
|
|
||||||
|
|
||||||
|
|
||||||
def load_analysis_json(path: Path) -> AnalysisJSON:
|
def load_analysis_json(path: Path) -> AnalysisJSON:
|
||||||
|
|||||||
@@ -363,3 +363,23 @@ class AnalysisJSON(TypedDict, total=False):
|
|||||||
completeness: list[AnalysisCompletenessRecord]
|
completeness: list[AnalysisCompletenessRecord]
|
||||||
duplicates: list[AnalysisDuplicateRecord]
|
duplicates: list[AnalysisDuplicateRecord]
|
||||||
|
|
||||||
|
|
||||||
|
class PlanOperationRecord(TypedDict, total=False):
|
||||||
|
operation_type: str
|
||||||
|
source_path: str
|
||||||
|
destination_path: str | None
|
||||||
|
reason: str
|
||||||
|
has_conflict: bool
|
||||||
|
conflict_reason: str | None
|
||||||
|
|
||||||
|
|
||||||
|
class PlanJSON(TypedDict, total=False):
|
||||||
|
vlm_schema_version: str
|
||||||
|
plan_id: str
|
||||||
|
created_at: str
|
||||||
|
operations: list[PlanOperationRecord]
|
||||||
|
summary: dict[str, int]
|
||||||
|
summary_by_reason: dict[str, int]
|
||||||
|
human_summary: str
|
||||||
|
metadata: dict[str, object]
|
||||||
|
|
||||||
|
|||||||
+106
-66
@@ -4,17 +4,24 @@ This module generates structured execution plans that specify how video files
|
|||||||
should be organized based on their parsed identities and configuration templates.
|
should be organized based on their parsed identities and configuration templates.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import json
|
|
||||||
import re
|
import re
|
||||||
import uuid
|
import uuid
|
||||||
from datetime import datetime
|
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Optional, Union
|
from typing import Optional, Union
|
||||||
|
|
||||||
from vlm.config import Config
|
from vlm.config import Config
|
||||||
from vlm.duplicate_resolve import choose_keep_index
|
from vlm.duplicate_resolve import DuplicateResolutionError, choose_keep_index
|
||||||
from vlm.io import validate_plan_json
|
from vlm.io import (
|
||||||
from vlm.utils import ensure_utc, is_sample_path, is_within_root, sanitize_path_component, utc_now
|
load_execution_plan,
|
||||||
|
save_execution_plan,
|
||||||
|
)
|
||||||
|
from vlm.utils import (
|
||||||
|
canonical_path_str,
|
||||||
|
is_sample_path,
|
||||||
|
is_within_root,
|
||||||
|
sanitize_path_component,
|
||||||
|
utc_now,
|
||||||
|
)
|
||||||
from vlm.models import (
|
from vlm.models import (
|
||||||
ExecutionPlan,
|
ExecutionPlan,
|
||||||
FileOperation,
|
FileOperation,
|
||||||
@@ -32,6 +39,42 @@ NO_OP_REASON_SEASON_OUT_OF_RANGE = "Series needs manual review (season exceeds c
|
|||||||
NO_OP_REASON_EPISODE_OUT_OF_RANGE = "Series needs manual review (episode exceeds configured threshold)"
|
NO_OP_REASON_EPISODE_OUT_OF_RANGE = "Series needs manual review (episode exceeds configured threshold)"
|
||||||
|
|
||||||
|
|
||||||
|
def _normalized_path_key(path_value: str | Path) -> str:
|
||||||
|
"""Normalize path-like values for duplicate-group matching."""
|
||||||
|
text = str(path_value).strip()
|
||||||
|
if not text:
|
||||||
|
return ""
|
||||||
|
return canonical_path_str(Path(text.replace("\\", "/")))
|
||||||
|
|
||||||
|
|
||||||
|
def _build_duplicate_quality_lookup(quality_comparison: list[dict]) -> dict[str, dict]:
|
||||||
|
"""Index duplicate quality entries by canonicalized path."""
|
||||||
|
lookup: dict[str, dict] = {}
|
||||||
|
for quality in quality_comparison:
|
||||||
|
quality_path = quality.get("path")
|
||||||
|
if isinstance(quality_path, str) and quality_path.strip():
|
||||||
|
lookup[_normalized_path_key(quality_path)] = quality
|
||||||
|
return lookup
|
||||||
|
|
||||||
|
|
||||||
|
def _mark_duplicate_group_manual_review(
|
||||||
|
operations: list[FileOperation],
|
||||||
|
indices: list[int],
|
||||||
|
message: str,
|
||||||
|
) -> None:
|
||||||
|
"""Convert unresolved duplicate operations into explicit manual-review no-ops."""
|
||||||
|
for index in indices:
|
||||||
|
current = operations[index]
|
||||||
|
operations[index] = FileOperation(
|
||||||
|
operation_type="no-op",
|
||||||
|
source_path=current.source_path,
|
||||||
|
destination_path=None,
|
||||||
|
reason=f"Duplicate group needs manual review: {message}",
|
||||||
|
has_conflict=False,
|
||||||
|
conflict_reason=None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def _analyze_directory_impact(operations: list[FileOperation]) -> dict:
|
def _analyze_directory_impact(operations: list[FileOperation]) -> dict:
|
||||||
"""Analyze which directories will be emptied by the plan."""
|
"""Analyze which directories will be emptied by the plan."""
|
||||||
# Get all source directories that have files being moved/renamed
|
# Get all source directories that have files being moved/renamed
|
||||||
@@ -99,17 +142,31 @@ def generate_plan(
|
|||||||
|
|
||||||
metadata: dict = {}
|
metadata: dict = {}
|
||||||
validation_snapshot: dict[str, object] = {}
|
validation_snapshot: dict[str, object] = {}
|
||||||
|
duplicate_resolution_issues: list[dict[str, object]] = []
|
||||||
if analysis_data is not None:
|
if analysis_data is not None:
|
||||||
validation_snapshot["analysis_source"] = analysis_data.get("metadata", {}).get("source_identities", "")
|
validation_snapshot["analysis_source"] = analysis_data.get("metadata", {}).get("source_identities", "")
|
||||||
validation_snapshot["duplicate_groups_considered"] = len(analysis_data.get("duplicates", []))
|
validation_snapshot["duplicate_groups_considered"] = len(analysis_data.get("duplicates", []))
|
||||||
validation_snapshot["completeness_seasons_with_gaps"] = len(analysis_data.get("completeness", []))
|
validation_snapshot["completeness_seasons_with_gaps"] = len(analysis_data.get("completeness", []))
|
||||||
|
|
||||||
if config.duplicate_keep != "manual":
|
if config.duplicate_keep != "manual":
|
||||||
path_to_index = {str(vf.path): i for i, (vf, _) in enumerate(identities)}
|
path_to_index = {
|
||||||
|
_normalized_path_key(vf.path): i for i, (vf, _) in enumerate(identities)
|
||||||
|
}
|
||||||
for dup in analysis_data.get("duplicates", []):
|
for dup in analysis_data.get("duplicates", []):
|
||||||
paths = dup.get("files", [])
|
paths = dup.get("files", [])
|
||||||
indices = [path_to_index[p] for p in paths if p in path_to_index]
|
normalized_paths = [
|
||||||
path_to_qc = {qc.get("path"): qc for qc in dup.get("quality_comparison", [])}
|
_normalized_path_key(path)
|
||||||
|
for path in paths
|
||||||
|
if isinstance(path, str) and path.strip()
|
||||||
|
]
|
||||||
|
indices = [
|
||||||
|
path_to_index[path_key]
|
||||||
|
for path_key in normalized_paths
|
||||||
|
if path_key in path_to_index
|
||||||
|
]
|
||||||
|
path_to_qc = _build_duplicate_quality_lookup(
|
||||||
|
dup.get("quality_comparison", [])
|
||||||
|
)
|
||||||
items = []
|
items = []
|
||||||
valid_indices = []
|
valid_indices = []
|
||||||
for i in indices:
|
for i in indices:
|
||||||
@@ -125,11 +182,25 @@ def generate_plan(
|
|||||||
if not items:
|
if not items:
|
||||||
continue
|
continue
|
||||||
quality_list = [
|
quality_list = [
|
||||||
path_to_qc.get(str(p), {}) for p, _ in items
|
path_to_qc.get(_normalized_path_key(path), {}) for path, _ in items
|
||||||
]
|
]
|
||||||
|
try:
|
||||||
|
if config.duplicate_keep == "by_quality" and any(not qc for qc in quality_list):
|
||||||
|
raise DuplicateResolutionError(
|
||||||
|
"missing quality comparison entries for one or more duplicate items"
|
||||||
|
)
|
||||||
keep_idx = choose_keep_index(
|
keep_idx = choose_keep_index(
|
||||||
items, config.duplicate_keep, quality_comparison=quality_list
|
items, config.duplicate_keep, quality_comparison=quality_list
|
||||||
)
|
)
|
||||||
|
except DuplicateResolutionError as exc:
|
||||||
|
issue = {
|
||||||
|
"strategy": config.duplicate_keep,
|
||||||
|
"reason": str(exc),
|
||||||
|
"files": [str(path) for path, _ in items],
|
||||||
|
}
|
||||||
|
duplicate_resolution_issues.append(issue)
|
||||||
|
_mark_duplicate_group_manual_review(operations, valid_indices, str(exc))
|
||||||
|
continue
|
||||||
if keep_idx is None:
|
if keep_idx is None:
|
||||||
continue
|
continue
|
||||||
keep_identity_index = valid_indices[keep_idx]
|
keep_identity_index = valid_indices[keep_idx]
|
||||||
@@ -165,14 +236,18 @@ def generate_plan(
|
|||||||
validation_snapshot["emptied_directories"] = [str(d) for d in directory_analysis["emptied_directories"]]
|
validation_snapshot["emptied_directories"] = [str(d) for d in directory_analysis["emptied_directories"]]
|
||||||
validation_snapshot["directory_warning"] = True
|
validation_snapshot["directory_warning"] = True
|
||||||
|
|
||||||
summary = _generate_summary(operations)
|
if duplicate_resolution_issues:
|
||||||
summary_by_reason = _generate_summary_by_reason(operations)
|
validation_snapshot["duplicate_resolution_issues"] = duplicate_resolution_issues
|
||||||
human_summary = _generate_human_summary(operations, summary, summary_by_reason, metadata)
|
metadata["duplicate_resolution_issues"] = duplicate_resolution_issues
|
||||||
|
|
||||||
if validation_snapshot:
|
if validation_snapshot:
|
||||||
validation_snapshot["captured_at"] = utc_now().isoformat()
|
validation_snapshot["captured_at"] = utc_now().isoformat()
|
||||||
metadata["validation_snapshot"] = validation_snapshot
|
metadata["validation_snapshot"] = validation_snapshot
|
||||||
|
|
||||||
|
summary = _generate_summary(operations)
|
||||||
|
summary_by_reason = _generate_summary_by_reason(operations)
|
||||||
|
human_summary = _generate_human_summary(operations, summary, summary_by_reason, metadata)
|
||||||
|
|
||||||
return ExecutionPlan(
|
return ExecutionPlan(
|
||||||
plan_id=str(uuid.uuid4()),
|
plan_id=str(uuid.uuid4()),
|
||||||
created_at=utc_now(),
|
created_at=utc_now(),
|
||||||
@@ -539,10 +614,25 @@ def _generate_human_summary(
|
|||||||
if conflicts > 0:
|
if conflicts > 0:
|
||||||
parts.append(f"冲突 {conflicts} 条。")
|
parts.append(f"冲突 {conflicts} 条。")
|
||||||
if metadata:
|
if metadata:
|
||||||
dup = metadata.get("duplicate_groups_considered", 0)
|
validation = metadata.get("validation_snapshot", {}) if isinstance(metadata.get("validation_snapshot"), dict) else {}
|
||||||
gaps = metadata.get("completeness_seasons_with_gaps", 0)
|
dup = metadata.get(
|
||||||
|
"duplicate_groups_considered",
|
||||||
|
validation.get("duplicate_groups_considered", 0),
|
||||||
|
)
|
||||||
|
gaps = metadata.get(
|
||||||
|
"completeness_seasons_with_gaps",
|
||||||
|
validation.get("completeness_seasons_with_gaps", 0),
|
||||||
|
)
|
||||||
if dup or gaps:
|
if dup or gaps:
|
||||||
parts.append(f"依据 analysis:重复组 {dup} 个;剧集缺口 {gaps} 季。")
|
parts.append(f"依据 analysis:重复组 {dup} 个;剧集缺口 {gaps} 季。")
|
||||||
|
resolution_issues = metadata.get(
|
||||||
|
"duplicate_resolution_issues",
|
||||||
|
validation.get("duplicate_resolution_issues", []),
|
||||||
|
)
|
||||||
|
if resolution_issues:
|
||||||
|
parts.append(
|
||||||
|
f"重复组中有 {len(resolution_issues)} 个因决策依据不足已转人工复核。"
|
||||||
|
)
|
||||||
quarantine_lines = _build_quarantine_recommendation_lines(operations)
|
quarantine_lines = _build_quarantine_recommendation_lines(operations)
|
||||||
if quarantine_lines:
|
if quarantine_lines:
|
||||||
parts.append("删除建议(仅隔离建议,执行删除前请人工复核):")
|
parts.append("删除建议(仅隔离建议,执行删除前请人工复核):")
|
||||||
@@ -598,31 +688,8 @@ def save_plan(plan: ExecutionPlan, output_path: Path) -> None:
|
|||||||
plan: ExecutionPlan to save
|
plan: ExecutionPlan to save
|
||||||
output_path: Path where the JSON file should be saved
|
output_path: Path where the JSON file should be saved
|
||||||
"""
|
"""
|
||||||
# Convert ExecutionPlan to dictionary
|
save_execution_plan(plan, output_path)
|
||||||
plan_dict = {
|
|
||||||
"vlm_schema_version": "1.0",
|
|
||||||
"plan_id": plan.plan_id,
|
|
||||||
"created_at": plan.created_at.isoformat(),
|
|
||||||
"operations": [
|
|
||||||
{
|
|
||||||
"operation_type": op.operation_type,
|
|
||||||
"source_path": str(op.source_path),
|
|
||||||
"destination_path": str(op.destination_path) if op.destination_path else None,
|
|
||||||
"reason": op.reason,
|
|
||||||
"has_conflict": op.has_conflict,
|
|
||||||
"conflict_reason": op.conflict_reason
|
|
||||||
}
|
|
||||||
for op in plan.operations
|
|
||||||
],
|
|
||||||
"summary": plan.summary,
|
|
||||||
"summary_by_reason": plan.summary_by_reason,
|
|
||||||
"human_summary": plan.human_summary,
|
|
||||||
"metadata": plan.metadata,
|
|
||||||
}
|
|
||||||
|
|
||||||
# Write to JSON file with indentation for human readability
|
|
||||||
with open(output_path, 'w', encoding='utf-8') as f:
|
|
||||||
json.dump(validate_plan_json(plan_dict), f, indent=2, ensure_ascii=False)
|
|
||||||
|
|
||||||
|
|
||||||
def load_plan(input_path: Path) -> ExecutionPlan:
|
def load_plan(input_path: Path) -> ExecutionPlan:
|
||||||
@@ -642,34 +709,7 @@ def load_plan(input_path: Path) -> ExecutionPlan:
|
|||||||
json.JSONDecodeError: If the file contains invalid JSON
|
json.JSONDecodeError: If the file contains invalid JSON
|
||||||
KeyError: If required fields are missing from the JSON
|
KeyError: If required fields are missing from the JSON
|
||||||
"""
|
"""
|
||||||
with open(input_path, 'r', encoding='utf-8') as f:
|
return load_execution_plan(input_path)
|
||||||
plan_dict = validate_plan_json(json.load(f))
|
|
||||||
|
|
||||||
# Reconstruct FileOperation objects
|
|
||||||
operations = [
|
|
||||||
FileOperation(
|
|
||||||
operation_type=op["operation_type"],
|
|
||||||
source_path=Path(op["source_path"]),
|
|
||||||
destination_path=Path(op["destination_path"]) if op["destination_path"] else None,
|
|
||||||
reason=op["reason"],
|
|
||||||
has_conflict=op["has_conflict"],
|
|
||||||
conflict_reason=op.get("conflict_reason")
|
|
||||||
)
|
|
||||||
for op in plan_dict["operations"]
|
|
||||||
]
|
|
||||||
|
|
||||||
# Reconstruct ExecutionPlan (normalize naive datetime to UTC for backward compatibility)
|
|
||||||
created_at = ensure_utc(datetime.fromisoformat(plan_dict["created_at"]))
|
|
||||||
return ExecutionPlan(
|
|
||||||
plan_id=plan_dict["plan_id"],
|
|
||||||
created_at=created_at,
|
|
||||||
operations=operations,
|
|
||||||
summary=plan_dict["summary"],
|
|
||||||
summary_by_reason=plan_dict.get("summary_by_reason", {}),
|
|
||||||
human_summary=plan_dict.get("human_summary", ""),
|
|
||||||
metadata=plan_dict.get("metadata", {}),
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def apply_review_to_plan(plan: ExecutionPlan, csv_path: Path) -> ExecutionPlan:
|
def apply_review_to_plan(plan: ExecutionPlan, csv_path: Path) -> ExecutionPlan:
|
||||||
"""Update a plan's operations based on a modified review CSV.
|
"""Update a plan's operations based on a modified review CSV.
|
||||||
|
|||||||
+12
-1
@@ -126,7 +126,18 @@ class QuarantineManager:
|
|||||||
operation_type="quarantine",
|
operation_type="quarantine",
|
||||||
file_path=file_path
|
file_path=file_path
|
||||||
)
|
)
|
||||||
raise ValueError(error_msg)
|
return OperationResult(
|
||||||
|
operation=FileOperation(
|
||||||
|
operation_type="quarantine",
|
||||||
|
source_path=file_path,
|
||||||
|
destination_path=None,
|
||||||
|
reason=reason or "Unsupported quarantine category",
|
||||||
|
has_conflict=False,
|
||||||
|
),
|
||||||
|
success=False,
|
||||||
|
error_message=error_msg,
|
||||||
|
executed_at=executed_at,
|
||||||
|
)
|
||||||
|
|
||||||
# Get the actual category directory name from the file path
|
# Get the actual category directory name from the file path
|
||||||
# (not the category name, which may differ due to category mappings)
|
# (not the category name, which may differ due to category mappings)
|
||||||
|
|||||||
+36
-20
@@ -5,14 +5,6 @@ from __future__ import annotations
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from textual import on
|
|
||||||
from textual.app import App, ComposeResult
|
|
||||||
from textual.binding import Binding
|
|
||||||
from textual.containers import Container, Horizontal, ScrollableContainer, Vertical
|
|
||||||
from textual.geometry import Size
|
|
||||||
from textual.screen import ModalScreen, Screen
|
|
||||||
from textual.widgets import DataTable, Footer, Static
|
|
||||||
|
|
||||||
from vlm.plan_review import save_review_csv
|
from vlm.plan_review import save_review_csv
|
||||||
from vlm.review_display import (
|
from vlm.review_display import (
|
||||||
build_csv_rows,
|
build_csv_rows,
|
||||||
@@ -21,14 +13,21 @@ from vlm.review_display import (
|
|||||||
risk_flags_to_labels,
|
risk_flags_to_labels,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
try:
|
||||||
|
from textual import on
|
||||||
|
from textual.app import App, ComposeResult
|
||||||
|
from textual.binding import Binding
|
||||||
|
from textual.containers import Container, Horizontal, ScrollableContainer
|
||||||
|
from textual.geometry import Size
|
||||||
|
from textual.screen import ModalScreen, Screen
|
||||||
|
from textual.widgets import DataTable, Footer, Static
|
||||||
|
except ImportError as exc: # pragma: no cover - exercised via runtime fallback
|
||||||
|
TEXTUAL_IMPORT_ERROR: ImportError | None = exc
|
||||||
|
else:
|
||||||
|
TEXTUAL_IMPORT_ERROR = None
|
||||||
|
|
||||||
def _index_from_row_key(row_key) -> int: # noqa: ANN001 - RowKey | str
|
|
||||||
if isinstance(row_key, str):
|
MISSING_TEXTUAL_MESSAGE = 'Textual is not installed. Install with: uv pip install -e ".[tui]"'
|
||||||
return int(row_key)
|
|
||||||
val = getattr(row_key, "value", None)
|
|
||||||
if val is None:
|
|
||||||
return int(row_key)
|
|
||||||
return int(val)
|
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
@@ -43,7 +42,18 @@ class ReviewTUIContext:
|
|||||||
summary_text: str
|
summary_text: str
|
||||||
|
|
||||||
|
|
||||||
class SummaryScreen(Screen):
|
if TEXTUAL_IMPORT_ERROR is None:
|
||||||
|
|
||||||
|
def _index_from_row_key(row_key) -> int: # noqa: ANN001 - RowKey | str
|
||||||
|
if isinstance(row_key, str):
|
||||||
|
return int(row_key)
|
||||||
|
val = getattr(row_key, "value", None)
|
||||||
|
if val is None:
|
||||||
|
return int(row_key)
|
||||||
|
return int(val)
|
||||||
|
|
||||||
|
|
||||||
|
class SummaryScreen(Screen):
|
||||||
"""Migration summary; Enter continues, q aborts."""
|
"""Migration summary; Enter continues, q aborts."""
|
||||||
|
|
||||||
BINDINGS = [
|
BINDINGS = [
|
||||||
@@ -84,7 +94,7 @@ class SummaryScreen(Screen):
|
|||||||
self.dismiss(False)
|
self.dismiss(False)
|
||||||
|
|
||||||
|
|
||||||
class ConfirmDiscardScreen(ModalScreen[bool]):
|
class ConfirmDiscardScreen(ModalScreen[bool]):
|
||||||
"""Confirm discarding unsaved edits."""
|
"""Confirm discarding unsaved edits."""
|
||||||
|
|
||||||
BINDINGS = [
|
BINDINGS = [
|
||||||
@@ -118,7 +128,7 @@ class ConfirmDiscardScreen(ModalScreen[bool]):
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
|
|
||||||
class ReviewMainScreen(Screen):
|
class ReviewMainScreen(Screen):
|
||||||
"""High-risk table + detail pane."""
|
"""High-risk table + detail pane."""
|
||||||
|
|
||||||
BINDINGS = [
|
BINDINGS = [
|
||||||
@@ -357,7 +367,7 @@ class ReviewMainScreen(Screen):
|
|||||||
self.app.push_screen(ConfirmDiscardScreen(), callback=after_confirm)
|
self.app.push_screen(ConfirmDiscardScreen(), callback=after_confirm)
|
||||||
|
|
||||||
|
|
||||||
class PlanReviewApp(App):
|
class PlanReviewApp(App):
|
||||||
"""Application shell: summary screen then review screen."""
|
"""Application shell: summary screen then review screen."""
|
||||||
|
|
||||||
def __init__(self, ctx: ReviewTUIContext) -> None:
|
def __init__(self, ctx: ReviewTUIContext) -> None:
|
||||||
@@ -380,9 +390,15 @@ class PlanReviewApp(App):
|
|||||||
self.exit(return_code=1)
|
self.exit(return_code=1)
|
||||||
|
|
||||||
|
|
||||||
def run_plan_review_tui(ctx: ReviewTUIContext) -> int:
|
def run_plan_review_tui(ctx: ReviewTUIContext) -> int:
|
||||||
"""Block until the user finishes the TUI. Returns process exit code."""
|
"""Block until the user finishes the TUI. Returns process exit code."""
|
||||||
app = PlanReviewApp(ctx)
|
app = PlanReviewApp(ctx)
|
||||||
app.run()
|
app.run()
|
||||||
code = app.return_code
|
code = app.return_code
|
||||||
return 0 if code is None else code
|
return 0 if code is None else code
|
||||||
|
|
||||||
|
else:
|
||||||
|
|
||||||
|
def run_plan_review_tui(ctx: ReviewTUIContext) -> int:
|
||||||
|
"""Raise a friendly error when the optional Textual dependency is missing."""
|
||||||
|
raise RuntimeError(MISSING_TEXTUAL_MESSAGE) from TEXTUAL_IMPORT_ERROR
|
||||||
|
|||||||
+27
-5
@@ -155,6 +155,29 @@ def _discover_video_paths(root: Path, video_extensions: list[str]) -> list[Path]
|
|||||||
return _discover_video_paths_recursive(root, video_extensions)
|
return _discover_video_paths_recursive(root, video_extensions)
|
||||||
|
|
||||||
|
|
||||||
|
def _log_find_nonzero_exit(returncode: int, stderr_text: str, discovered_count: int) -> None:
|
||||||
|
"""Log the explicit contract for non-zero `find` exits.
|
||||||
|
|
||||||
|
Contract: if `find` emits partial stdout before failing, keep those paths and
|
||||||
|
continue with a warning. If no paths were emitted, return an empty result and
|
||||||
|
log that scan discovery was incomplete.
|
||||||
|
"""
|
||||||
|
stderr_suffix = f": {stderr_text}" if stderr_text else ""
|
||||||
|
if discovered_count > 0:
|
||||||
|
logger.warning(
|
||||||
|
"find exited with code %s; using %s partial scan result(s)%s",
|
||||||
|
returncode,
|
||||||
|
discovered_count,
|
||||||
|
stderr_suffix,
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
logger.warning(
|
||||||
|
"find exited with code %s and produced no scan results%s",
|
||||||
|
returncode,
|
||||||
|
stderr_suffix,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def _discover_video_paths_with_find(root: Path, video_extensions: list[str]) -> list[Path]:
|
def _discover_video_paths_with_find(root: Path, video_extensions: list[str]) -> list[Path]:
|
||||||
"""Discover matching video files using the system `find` command."""
|
"""Discover matching video files using the system `find` command."""
|
||||||
normalized_extensions = [ext.lower() for ext in video_extensions if ext]
|
normalized_extensions = [ext.lower() for ext in video_extensions if ext]
|
||||||
@@ -175,11 +198,6 @@ def _discover_video_paths_with_find(root: Path, video_extensions: list[str]) ->
|
|||||||
)
|
)
|
||||||
stdout, stderr = process.communicate()
|
stdout, stderr = process.communicate()
|
||||||
|
|
||||||
if process.returncode != 0:
|
|
||||||
stderr_text = stderr.decode(errors="replace").strip()
|
|
||||||
if stderr_text:
|
|
||||||
logger.warning(f"find reported issues while scanning: {stderr_text}")
|
|
||||||
|
|
||||||
discovered_paths: list[Path] = []
|
discovered_paths: list[Path] = []
|
||||||
for path_bytes in stdout.split(b"\0"):
|
for path_bytes in stdout.split(b"\0"):
|
||||||
if not path_bytes:
|
if not path_bytes:
|
||||||
@@ -189,6 +207,10 @@ def _discover_video_paths_with_find(root: Path, video_extensions: list[str]) ->
|
|||||||
continue
|
continue
|
||||||
discovered_paths.append(file_path)
|
discovered_paths.append(file_path)
|
||||||
|
|
||||||
|
if process.returncode != 0:
|
||||||
|
stderr_text = stderr.decode(errors="replace").strip()
|
||||||
|
_log_find_nonzero_exit(process.returncode, stderr_text, len(discovered_paths))
|
||||||
|
|
||||||
return discovered_paths
|
return discovered_paths
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -4,7 +4,7 @@ import pytest
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from vlm.models import MovieIdentity
|
from vlm.models import MovieIdentity
|
||||||
from vlm.duplicate_resolve import choose_keep_index
|
from vlm.duplicate_resolve import DuplicateResolutionError, choose_keep_index
|
||||||
|
|
||||||
|
|
||||||
def _mi(title: str = "Test", year: int | None = 2020) -> MovieIdentity:
|
def _mi(title: str = "Test", year: int | None = 2020) -> MovieIdentity:
|
||||||
@@ -218,3 +218,21 @@ class TestOtherStrategies:
|
|||||||
]
|
]
|
||||||
idx = choose_keep_index(items, "by_reputation_quality_time", quality_comparison=qc)
|
idx = choose_keep_index(items, "by_reputation_quality_time", quality_comparison=qc)
|
||||||
assert idx == 1
|
assert idx == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_by_quality_requires_aligned_quality_data():
|
||||||
|
items = [(Path("/a.mkv"), _mi()), (Path("/b.mkv"), _mi())]
|
||||||
|
|
||||||
|
with pytest.raises(
|
||||||
|
DuplicateResolutionError,
|
||||||
|
match="requires quality data aligned with duplicate items",
|
||||||
|
):
|
||||||
|
choose_keep_index(items, "by_quality", quality_comparison=[{"path": "/a.mkv"}])
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_strategy_raises_duplicate_resolution_error():
|
||||||
|
items = [(Path("/a.mkv"), _mi()), (Path("/b.mkv"), _mi())]
|
||||||
|
|
||||||
|
with pytest.raises(DuplicateResolutionError, match="Unsupported duplicate strategy"):
|
||||||
|
choose_keep_index(items, "unexpected")
|
||||||
|
|
||||||
|
|||||||
@@ -946,3 +946,59 @@ class TestRollbackExecution:
|
|||||||
assert original_path.exists()
|
assert original_path.exists()
|
||||||
assert original_path.read_text() == "movie content"
|
assert original_path.read_text() == "movie content"
|
||||||
assert not quarantine_path.exists()
|
assert not quarantine_path.exists()
|
||||||
|
|
||||||
|
|
||||||
|
def test_execute_plan_continues_after_unexpected_operation_exception(temp_test_dir):
|
||||||
|
logger = logging.getLogger("test_executor")
|
||||||
|
logger.setLevel(logging.DEBUG)
|
||||||
|
engine = ExecutionEngine(logger=logger)
|
||||||
|
|
||||||
|
source_ok = temp_test_dir["test_file1"]
|
||||||
|
destination_ok = temp_test_dir["dest_dir"] / "moved1.mp4"
|
||||||
|
source_broken = temp_test_dir["test_file2"]
|
||||||
|
destination_broken = temp_test_dir["dest_dir"] / "broken.mkv"
|
||||||
|
|
||||||
|
operations = [
|
||||||
|
FileOperation(
|
||||||
|
operation_type="move",
|
||||||
|
source_path=source_broken,
|
||||||
|
destination_path=destination_broken,
|
||||||
|
reason="broken operation",
|
||||||
|
has_conflict=False,
|
||||||
|
conflict_reason=None,
|
||||||
|
),
|
||||||
|
FileOperation(
|
||||||
|
operation_type="move",
|
||||||
|
source_path=source_ok,
|
||||||
|
destination_path=destination_ok,
|
||||||
|
reason="healthy operation",
|
||||||
|
has_conflict=False,
|
||||||
|
conflict_reason=None,
|
||||||
|
),
|
||||||
|
]
|
||||||
|
plan = ExecutionPlan(
|
||||||
|
plan_id=str(uuid4()),
|
||||||
|
created_at=datetime.now(timezone.utc),
|
||||||
|
operations=operations,
|
||||||
|
summary={"move": 2},
|
||||||
|
)
|
||||||
|
|
||||||
|
original_execute_operation = engine.execute_operation
|
||||||
|
|
||||||
|
def flaky_execute_operation(operation, mode):
|
||||||
|
if operation.source_path == source_broken:
|
||||||
|
raise RuntimeError("boom")
|
||||||
|
return original_execute_operation(operation, mode)
|
||||||
|
|
||||||
|
engine.execute_operation = flaky_execute_operation # type: ignore[method-assign]
|
||||||
|
|
||||||
|
results, summary, _ = engine.execute_plan(plan, mode="execute", confirmed=True)
|
||||||
|
|
||||||
|
assert len(results) == 2
|
||||||
|
assert results[0].success is False
|
||||||
|
assert "Unexpected failure during move: boom" == results[0].error_message
|
||||||
|
assert results[1].success is True
|
||||||
|
assert summary["failed"] == 1
|
||||||
|
assert summary["successful"] == 1
|
||||||
|
assert source_broken.exists()
|
||||||
|
assert destination_ok.exists()
|
||||||
|
|||||||
+99
-1
@@ -8,11 +8,15 @@ import pytest
|
|||||||
|
|
||||||
from vlm.io import (
|
from vlm.io import (
|
||||||
_video_file_from_record,
|
_video_file_from_record,
|
||||||
|
execution_plan_from_record,
|
||||||
|
execution_plan_to_record,
|
||||||
load_analysis_json,
|
load_analysis_json,
|
||||||
|
load_execution_plan,
|
||||||
load_identities_json,
|
load_identities_json,
|
||||||
|
save_execution_plan,
|
||||||
save_identities_json,
|
save_identities_json,
|
||||||
)
|
)
|
||||||
from vlm.models import VideoFile
|
from vlm.models import ExecutionPlan, FileOperation, VideoFile
|
||||||
|
|
||||||
|
|
||||||
class TestVideoFileFromRecord:
|
class TestVideoFileFromRecord:
|
||||||
@@ -105,6 +109,100 @@ class TestVideoFileFromRecord:
|
|||||||
assert vf.bitrate_kbps is None
|
assert vf.bitrate_kbps is None
|
||||||
|
|
||||||
|
|
||||||
|
class TestExecutionPlanIo:
|
||||||
|
"""Tests for the validated typed execution plan boundary."""
|
||||||
|
|
||||||
|
def test_execution_plan_record_round_trip(self):
|
||||||
|
plan = ExecutionPlan(
|
||||||
|
plan_id="plan-123",
|
||||||
|
created_at=datetime(2026, 4, 7, 12, 0, tzinfo=timezone.utc),
|
||||||
|
operations=[
|
||||||
|
FileOperation(
|
||||||
|
operation_type="move",
|
||||||
|
source_path=Path("/library/movie/source.mkv"),
|
||||||
|
destination_path=Path("/library/movie/target.mkv"),
|
||||||
|
reason="move movie",
|
||||||
|
has_conflict=False,
|
||||||
|
conflict_reason=None,
|
||||||
|
)
|
||||||
|
],
|
||||||
|
summary={"total": 1, "move": 1, "rename": 0, "quarantine": 0, "no-op": 0},
|
||||||
|
summary_by_reason={"move movie": 1},
|
||||||
|
human_summary="计划已生成",
|
||||||
|
metadata={"analysis_source": "analysis.json"},
|
||||||
|
)
|
||||||
|
|
||||||
|
record = execution_plan_to_record(plan)
|
||||||
|
loaded = execution_plan_from_record(record)
|
||||||
|
|
||||||
|
assert loaded.plan_id == plan.plan_id
|
||||||
|
assert loaded.created_at == plan.created_at
|
||||||
|
assert loaded.operations[0].source_path == plan.operations[0].source_path
|
||||||
|
assert loaded.operations[0].destination_path == plan.operations[0].destination_path
|
||||||
|
assert loaded.summary == plan.summary
|
||||||
|
assert loaded.summary_by_reason == plan.summary_by_reason
|
||||||
|
assert loaded.human_summary == plan.human_summary
|
||||||
|
assert loaded.metadata == plan.metadata
|
||||||
|
|
||||||
|
def test_load_execution_plan_normalizes_naive_timestamp(self, tmp_path):
|
||||||
|
plan_path = tmp_path / "plan.json"
|
||||||
|
plan_path.write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"vlm_schema_version": "1.0",
|
||||||
|
"plan_id": "plan-naive",
|
||||||
|
"created_at": "2026-04-07T12:00:00",
|
||||||
|
"operations": [
|
||||||
|
{
|
||||||
|
"operation_type": "no-op",
|
||||||
|
"source_path": "/library/movie/source.mkv",
|
||||||
|
"destination_path": None,
|
||||||
|
"reason": "manual review",
|
||||||
|
"has_conflict": False,
|
||||||
|
"conflict_reason": None,
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {"total": 1, "move": 0, "rename": 0, "quarantine": 0, "no-op": 1},
|
||||||
|
"summary_by_reason": {"manual review": 1},
|
||||||
|
"human_summary": "summary",
|
||||||
|
"metadata": {},
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
|
||||||
|
loaded = load_execution_plan(plan_path)
|
||||||
|
|
||||||
|
assert loaded.created_at.tzinfo == timezone.utc
|
||||||
|
assert loaded.created_at.isoformat() == "2026-04-07T12:00:00+00:00"
|
||||||
|
|
||||||
|
def test_save_execution_plan_writes_validated_schema(self, tmp_path):
|
||||||
|
plan = ExecutionPlan(
|
||||||
|
plan_id="plan-save",
|
||||||
|
created_at=datetime(2026, 4, 7, 13, 0, tzinfo=timezone.utc),
|
||||||
|
operations=[
|
||||||
|
FileOperation(
|
||||||
|
operation_type="quarantine",
|
||||||
|
source_path=Path("/library/movie/duplicate.mkv"),
|
||||||
|
destination_path=None,
|
||||||
|
reason="duplicate",
|
||||||
|
has_conflict=False,
|
||||||
|
conflict_reason=None,
|
||||||
|
)
|
||||||
|
],
|
||||||
|
summary={"total": 1, "move": 0, "rename": 0, "quarantine": 1, "no-op": 0},
|
||||||
|
)
|
||||||
|
|
||||||
|
output_path = tmp_path / "plan.json"
|
||||||
|
save_execution_plan(plan, output_path)
|
||||||
|
saved = json.loads(output_path.read_text(encoding="utf-8"))
|
||||||
|
|
||||||
|
assert saved["vlm_schema_version"] == "1.0"
|
||||||
|
assert saved["plan_id"] == "plan-save"
|
||||||
|
assert saved["operations"][0]["operation_type"] == "quarantine"
|
||||||
|
assert saved["operations"][0]["source_path"] == "/library/movie/duplicate.mkv"
|
||||||
|
|
||||||
|
|
||||||
class TestIdentitiesJsonVersioning:
|
class TestIdentitiesJsonVersioning:
|
||||||
"""Tests for identities.json schema versioning."""
|
"""Tests for identities.json schema versioning."""
|
||||||
|
|
||||||
|
|||||||
@@ -101,3 +101,37 @@ def test_executor_blocks_unsafe_destination_even_with_manual_plan(tmp_path):
|
|||||||
assert not results[0].success
|
assert not results[0].success
|
||||||
assert "outside library root" in (results[0].error_message or "")
|
assert "outside library root" in (results[0].error_message or "")
|
||||||
assert summary["failed"] == 1
|
assert summary["failed"] == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_executor_blocks_unsafe_source_even_with_manual_plan(tmp_path):
|
||||||
|
library_root = tmp_path / "library"
|
||||||
|
library_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
outside_root = tmp_path / "outside"
|
||||||
|
outside_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
source = outside_root / "Sample.mkv"
|
||||||
|
source.write_text("sample")
|
||||||
|
|
||||||
|
destination = library_root / "movie" / "Sample.mkv"
|
||||||
|
operation = FileOperation(
|
||||||
|
operation_type="move",
|
||||||
|
source_path=source,
|
||||||
|
destination_path=destination,
|
||||||
|
reason="unsafe test",
|
||||||
|
has_conflict=False,
|
||||||
|
conflict_reason=None,
|
||||||
|
)
|
||||||
|
plan = ExecutionPlan(
|
||||||
|
plan_id=str(uuid4()),
|
||||||
|
created_at=datetime.now(timezone.utc),
|
||||||
|
operations=[operation],
|
||||||
|
summary={"total": 1, "move": 1, "rename": 0, "quarantine": 0, "no-op": 0},
|
||||||
|
)
|
||||||
|
|
||||||
|
engine = ExecutionEngine(config=Config(library_root=library_root))
|
||||||
|
results, summary, _ = engine.execute_plan(plan, mode="execute", confirmed=True)
|
||||||
|
|
||||||
|
assert not results[0].success
|
||||||
|
assert "Unsafe source outside library root" in (results[0].error_message or "")
|
||||||
|
assert summary["failed"] == 1
|
||||||
|
assert source.exists()
|
||||||
|
assert not destination.exists()
|
||||||
|
|||||||
@@ -614,6 +614,114 @@ def test_generate_plan_with_analysis_by_reputation_missing_scores_uses_fallback_
|
|||||||
assert "评分依据不足" in plan.human_summary
|
assert "评分依据不足" in plan.human_summary
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
def test_generate_plan_with_analysis_by_quality_missing_quality_data_marks_manual_review(config):
|
||||||
|
"""Missing quality entries should not silently keep the first duplicate."""
|
||||||
|
config.duplicate_keep = "by_quality"
|
||||||
|
p_low = Path("/mnt/nas/videos/movie/Test.2020.720p.WEB-DL.mkv")
|
||||||
|
p_high = Path("/mnt/nas/videos/movie/Test.2020.1080p.BluRay.mkv")
|
||||||
|
vf_low = VideoFile(
|
||||||
|
path=p_low,
|
||||||
|
filename="Test.2020.720p.WEB-DL.mkv",
|
||||||
|
size_bytes=1000000,
|
||||||
|
modified_timestamp=datetime.now(timezone.utc),
|
||||||
|
category="movie",
|
||||||
|
)
|
||||||
|
vf_high = VideoFile(
|
||||||
|
path=p_high,
|
||||||
|
filename="Test.2020.1080p.BluRay.mkv",
|
||||||
|
size_bytes=2000000,
|
||||||
|
modified_timestamp=datetime.now(timezone.utc),
|
||||||
|
category="movie",
|
||||||
|
)
|
||||||
|
identity = MovieIdentity(
|
||||||
|
title="Test",
|
||||||
|
year=2020,
|
||||||
|
confidence=0.9,
|
||||||
|
needs_review=False,
|
||||||
|
original_filename="Test.2020.mkv",
|
||||||
|
)
|
||||||
|
|
||||||
|
plan = generate_plan(
|
||||||
|
[(vf_low, identity), (vf_high, identity)],
|
||||||
|
config,
|
||||||
|
analysis_data={
|
||||||
|
"metadata": {"source_identities": "identities.json"},
|
||||||
|
"completeness": [],
|
||||||
|
"duplicates": [
|
||||||
|
{
|
||||||
|
"identity": {"type": "movie", "title": "Test", "year": 2020},
|
||||||
|
"files": [str(p_low), str(p_high)],
|
||||||
|
"quality_comparison": [
|
||||||
|
{"path": str(p_low), "resolution": "1280x720", "size_bytes": 1000000},
|
||||||
|
],
|
||||||
|
}
|
||||||
|
],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
assert all(op.operation_type == "no-op" for op in plan.operations)
|
||||||
|
assert "Duplicate group needs manual review" in plan.operations[0].reason
|
||||||
|
issues = plan.metadata["validation_snapshot"]["duplicate_resolution_issues"]
|
||||||
|
assert len(issues) == 1
|
||||||
|
assert "missing quality comparison entries" in issues[0]["reason"]
|
||||||
|
assert "转人工复核" in plan.human_summary
|
||||||
|
|
||||||
|
|
||||||
|
def test_generate_plan_normalizes_duplicate_paths_before_matching(config):
|
||||||
|
"""Duplicate matching should survive harmless path-format differences."""
|
||||||
|
config.duplicate_keep = "by_quality"
|
||||||
|
p_720 = Path("/mnt/nas/videos/movie/Test.2020.720p.WEB-DL.mkv")
|
||||||
|
p_1080 = Path("/mnt/nas/videos/movie/Test.2020.1080p.BluRay.mkv")
|
||||||
|
vf_720 = VideoFile(
|
||||||
|
path=p_720,
|
||||||
|
filename="Test.2020.720p.WEB-DL.mkv",
|
||||||
|
size_bytes=1000000,
|
||||||
|
modified_timestamp=datetime.now(timezone.utc),
|
||||||
|
category="movie",
|
||||||
|
)
|
||||||
|
vf_1080 = VideoFile(
|
||||||
|
path=p_1080,
|
||||||
|
filename="Test.2020.1080p.BluRay.mkv",
|
||||||
|
size_bytes=2000000,
|
||||||
|
modified_timestamp=datetime.now(timezone.utc),
|
||||||
|
category="movie",
|
||||||
|
)
|
||||||
|
identity = MovieIdentity(
|
||||||
|
title="Test",
|
||||||
|
year=2020,
|
||||||
|
confidence=0.9,
|
||||||
|
needs_review=False,
|
||||||
|
original_filename="Test.2020.mkv",
|
||||||
|
)
|
||||||
|
|
||||||
|
normalized_low = str(p_720.parent / "." / p_720.name)
|
||||||
|
normalized_high = str(p_1080.parent / "." / p_1080.name)
|
||||||
|
plan = generate_plan(
|
||||||
|
[(vf_720, identity), (vf_1080, identity)],
|
||||||
|
config,
|
||||||
|
analysis_data={
|
||||||
|
"metadata": {"source_identities": "identities.json"},
|
||||||
|
"completeness": [],
|
||||||
|
"duplicates": [
|
||||||
|
{
|
||||||
|
"identity": {"type": "movie", "title": "Test", "year": 2020},
|
||||||
|
"files": [normalized_low, normalized_high],
|
||||||
|
"quality_comparison": [
|
||||||
|
{"path": normalized_low, "resolution": "1280x720", "size_bytes": 1000000},
|
||||||
|
{"path": normalized_high, "resolution": "1920x1080", "size_bytes": 2000000},
|
||||||
|
],
|
||||||
|
}
|
||||||
|
],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
assert plan.summary["quarantine"] == 1
|
||||||
|
assert plan.operations[0].operation_type == "quarantine"
|
||||||
|
assert plan.operations[1].operation_type == "move"
|
||||||
|
|
||||||
|
|
||||||
def test_generate_plan_human_summary_marks_disc_files_as_high_risk(config):
|
def test_generate_plan_human_summary_marks_disc_files_as_high_risk(config):
|
||||||
"""Disc/part files should be listed with a conservative risk note in summary."""
|
"""Disc/part files should be listed with a conservative risk note in summary."""
|
||||||
config.duplicate_keep = "by_reputation"
|
config.duplicate_keep = "by_reputation"
|
||||||
|
|||||||
@@ -82,27 +82,29 @@ class TestQuarantineManager:
|
|||||||
assert expected_quarantine_path.read_text() == "test content"
|
assert expected_quarantine_path.read_text() == "test content"
|
||||||
|
|
||||||
def test_quarantine_anime_file_rejected(self, manager, config):
|
def test_quarantine_anime_file_rejected(self, manager, config):
|
||||||
"""Test that quarantining anime files is rejected."""
|
"""Test that quarantining anime files returns a failed result."""
|
||||||
# Create a test anime file
|
# Create a test anime file
|
||||||
anime_file = config.library_root / "anime" / "Test Anime.mkv"
|
anime_file = config.library_root / "anime" / "Test Anime.mkv"
|
||||||
anime_file.write_text("test content")
|
anime_file.write_text("test content")
|
||||||
|
|
||||||
# Attempt to quarantine should raise ValueError
|
result = manager.quarantine_file(anime_file)
|
||||||
with pytest.raises(ValueError, match="Quarantine not supported for category 'anime'"):
|
|
||||||
manager.quarantine_file(anime_file)
|
assert result.success is False
|
||||||
|
assert "Quarantine not supported for category 'anime'" in (result.error_message or "")
|
||||||
|
|
||||||
# Verify file was not moved
|
# Verify file was not moved
|
||||||
assert anime_file.exists()
|
assert anime_file.exists()
|
||||||
|
|
||||||
def test_quarantine_other_file_rejected(self, manager, config):
|
def test_quarantine_other_file_rejected(self, manager, config):
|
||||||
"""Test that quarantining other files is rejected."""
|
"""Test that quarantining other files returns a failed result."""
|
||||||
# Create a test other file
|
# Create a test other file
|
||||||
other_file = config.library_root / "other" / "Test File.mkv"
|
other_file = config.library_root / "other" / "Test File.mkv"
|
||||||
other_file.write_text("test content")
|
other_file.write_text("test content")
|
||||||
|
|
||||||
# Attempt to quarantine should raise ValueError
|
result = manager.quarantine_file(other_file)
|
||||||
with pytest.raises(ValueError, match="Quarantine not supported for category 'other'"):
|
|
||||||
manager.quarantine_file(other_file)
|
assert result.success is False
|
||||||
|
assert "Quarantine not supported for category 'other'" in (result.error_message or "")
|
||||||
|
|
||||||
# Verify file was not moved
|
# Verify file was not moved
|
||||||
assert other_file.exists()
|
assert other_file.exists()
|
||||||
|
|||||||
@@ -137,6 +137,54 @@ class TestScanLibrary:
|
|||||||
assert len(result) == 1
|
assert len(result) == 1
|
||||||
assert result[0].path == visible_file
|
assert result[0].path == visible_file
|
||||||
|
|
||||||
|
def test_scan_keeps_partial_find_results_when_find_exits_nonzero(self, tmp_path):
|
||||||
|
"""Non-zero find exits should keep partial stdout and log the contract."""
|
||||||
|
movie_dir = tmp_path / "movie"
|
||||||
|
movie_dir.mkdir()
|
||||||
|
visible_file = movie_dir / "visible.mp4"
|
||||||
|
visible_file.touch()
|
||||||
|
|
||||||
|
fake_stdout = f"{visible_file}\0".encode()
|
||||||
|
|
||||||
|
with patch('subprocess.Popen') as mock_popen, patch("vlm.scanner.logger.warning") as mock_warning:
|
||||||
|
process = MagicMock()
|
||||||
|
process.communicate.return_value = (fake_stdout, b"Permission denied")
|
||||||
|
process.returncode = 1
|
||||||
|
mock_popen.return_value = process
|
||||||
|
|
||||||
|
config = Config(library_root=tmp_path)
|
||||||
|
result = scan_library(tmp_path, config, include_video_metadata=False)
|
||||||
|
|
||||||
|
warning_messages = [
|
||||||
|
call.args[0] % call.args[1:] if len(call.args) > 1 else call.args[0]
|
||||||
|
for call in mock_warning.call_args_list
|
||||||
|
]
|
||||||
|
assert len(result) == 1
|
||||||
|
assert result[0].path == visible_file
|
||||||
|
assert any("using 1 partial scan result" in message for message in warning_messages)
|
||||||
|
assert any("Permission denied" in message for message in warning_messages)
|
||||||
|
|
||||||
|
def test_scan_returns_empty_when_find_exits_nonzero_without_stdout(self, tmp_path):
|
||||||
|
"""Non-zero find exits without stdout should produce an empty result deterministically."""
|
||||||
|
(tmp_path / "movie").mkdir()
|
||||||
|
|
||||||
|
with patch('subprocess.Popen') as mock_popen, patch("vlm.scanner.logger.warning") as mock_warning:
|
||||||
|
process = MagicMock()
|
||||||
|
process.communicate.return_value = (b"", b"Permission denied")
|
||||||
|
process.returncode = 1
|
||||||
|
mock_popen.return_value = process
|
||||||
|
|
||||||
|
config = Config(library_root=tmp_path)
|
||||||
|
result = scan_library(tmp_path, config, include_video_metadata=False)
|
||||||
|
|
||||||
|
warning_messages = [
|
||||||
|
call.args[0] % call.args[1:] if len(call.args) > 1 else call.args[0]
|
||||||
|
for call in mock_warning.call_args_list
|
||||||
|
]
|
||||||
|
assert result == []
|
||||||
|
assert any("produced no scan results" in message for message in warning_messages)
|
||||||
|
assert any("Permission denied" in message for message in warning_messages)
|
||||||
|
|
||||||
def test_scan_falls_back_when_find_is_unavailable(self, tmp_path):
|
def test_scan_falls_back_when_find_is_unavailable(self, tmp_path):
|
||||||
"""Test scan_library falls back to recursive scanning if find is unavailable."""
|
"""Test scan_library falls back to recursive scanning if find is unavailable."""
|
||||||
movie_dir = tmp_path / "movie"
|
movie_dir = tmp_path / "movie"
|
||||||
|
|||||||
Reference in New Issue
Block a user