From 5931f16a71987d73a350f7f0639b5ec2fe88f2ba Mon Sep 17 00:00:00 2001 From: windyboy Date: Fri, 13 Feb 2026 12:01:10 +0800 Subject: [PATCH] merge VLM skills into a single workflow expert skill --- skills/vlm-library-workflow/SKILL.md | 94 +++++++++++++++++++ .../vlm-library-workflow/agents/openai.yaml | 4 + .../references/cli-reference.md | 26 +++++ .../references/command-recipes.md | 71 ++++++++++++++ .../references/dev-guide.md | 30 ++++++ .../references/workflow.md | 67 +++++++++++++ 6 files changed, 292 insertions(+) create mode 100644 skills/vlm-library-workflow/SKILL.md create mode 100644 skills/vlm-library-workflow/agents/openai.yaml create mode 100644 skills/vlm-library-workflow/references/cli-reference.md create mode 100644 skills/vlm-library-workflow/references/command-recipes.md create mode 100644 skills/vlm-library-workflow/references/dev-guide.md create mode 100644 skills/vlm-library-workflow/references/workflow.md diff --git a/skills/vlm-library-workflow/SKILL.md b/skills/vlm-library-workflow/SKILL.md new file mode 100644 index 0000000..703624c --- /dev/null +++ b/skills/vlm-library-workflow/SKILL.md @@ -0,0 +1,94 @@ +--- +name: vlm-library-workflow +description: Operate and extend the Video Library Manager (`vlm`) with a safety-first, human-in-the-loop workflow across scan, parse, enrich, analyze, plan, review-plan, execute, rollback, and developer verification. Use when requests involve organizing a video library, producing or reviewing `inventory.csv`/`identities.json`/`analysis.json`/`plan.json`, tuning VLM config templates, resolving duplicates or episode gaps, running dry-run/confirm execution, recovering changes via rollback, explaining VLM CLI usage, or developing/modifying VLM features (parser, providers, planner, executor, commands, tests). +--- + +# VLM Library Workflow + +## Overview + +Run the repository's built-in media organization pipeline with consistent safety checks and explicit output verification. Prefer incremental, reviewable steps and never skip dry-run and plan inspection before destructive operations. + +## Workflow Order + +Run commands in this default sequence unless the user asks for a specific stage: + +1. `vlm config init` then set `library_root` in `~/.vlm/config.yaml` +2. `vlm scan` to produce `inventory.csv` +3. `vlm parse --inventory inventory.csv` to produce metadata-rich `identities.json` +4. `vlm enrich` when bilingual titles and reputation signals are needed +5. `vlm analyze --inventory inventory.csv` to produce `analysis.json` +6. `vlm plan --analysis analysis.json` to produce `plan.json` +7. `vlm execute` (dry-run) and inspect summary +8. `vlm execute --confirm` only after explicit user confirmation +9. `vlm rollback` if the user requests revert + +## Preflight Checks + +Run these checks before executing workflow commands: + +1. Confirm current working directory is repository root. +2. Probe command availability in this order: + 1. `vlm --help` + 2. if unavailable, switch all commands to `uv run vlm --help` +3. Run the target subcommand `--help` when options are uncertain. +4. Confirm config validity with `vlm config validate` after config edits. +5. Verify required input files exist before downstream stages. +6. Treat `execute --confirm` as destructive and require explicit user confirmation. + +## Execution Rules + +Follow these rules while executing tasks: + +1. Prefer read-only stages first: scan, parse, enrich, analyze, plan. +2. Treat `plan.json` as reviewable contract; summarize counts and conflicts before execution. +3. Run dry-run (`vlm execute`) before `vlm execute --confirm`. +4. If execution is interrupted or results are incorrect, locate rollback logs and run `vlm rollback`. +5. Keep outputs explicit in responses: file path, record counts, and next command. +6. If user asks for partial workflow, run only required stages and clearly state skipped dependencies. +7. Before `execute --confirm`, run `vlm review-plan --input plan.json --output plan_manual_review.csv` and report high-risk counts. +8. If high-risk operations > 0, default to pause and require user explicit override to continue confirm execution. + +## Decision Points + +Use these decision policies: + +1. Duplicate handling: + Choose `plan.duplicate_keep` policy (`by_quality`, `by_reputation`, `first_seen`, `manual`) based on user preference. +2. Enrichment: + Skip `vlm enrich` only when user does not need translations/reputation or API keys are unavailable. +3. Metadata quality: + Prefer `vlm parse --inventory inventory.csv` when duplicate quality ranking matters. +4. Analysis-assisted planning: + Prefer `vlm plan --analysis analysis.json` when user wants automatic duplicate quarantine decisions. +5. Parser boundary risk: + Treat filenames containing resolution-like `1920x1080/1440x1080` and `Sample` clips as high-risk; require review-plan output before confirmation. + +## Output Contract + +Return concise, operational summaries: + +1. Commands executed. +2. Artifacts generated or updated. +3. Key counts (files scanned, identities parsed, duplicate groups, plan operations). +4. Risk counts from `review-plan` (manual_review / sample_source / high_season / high_episode / conflicts). +5. Blocking errors and exact remediation command. +6. Safe next step. + +## Developer Verification + +When modifying VLM core logic, verify integrity with these recipes: + +1. **Full Suite**: `pytest` +2. **Core Components**: `pytest tests/test_scanner.py tests/test_planner.py tests/test_executor.py` +3. **Property Tests**: `pytest tests/test_analysis_properties.py` +4. **Integration**: `pytest tests/test_reports_integration.py` + +## References + +Load these references on demand: + +1. `references/command-recipes.md`: command syntax, artifact expectations, and failure triage. +2. `references/workflow.md`: end-to-end operator workflow guidance. +3. `references/cli-reference.md`: command and config option quick reference. +4. `references/dev-guide.md`: project architecture and developer modification patterns. diff --git a/skills/vlm-library-workflow/agents/openai.yaml b/skills/vlm-library-workflow/agents/openai.yaml new file mode 100644 index 0000000..45ece66 --- /dev/null +++ b/skills/vlm-library-workflow/agents/openai.yaml @@ -0,0 +1,4 @@ +interface: + display_name: "VLM Workflow Expert" + short_description: "Run and extend VLM safely" + default_prompt: "Use this skill for safety-first VLM operations and development: scan, parse, enrich, analyze, plan, review-plan, execute, rollback, and targeted code/test updates." diff --git a/skills/vlm-library-workflow/references/cli-reference.md b/skills/vlm-library-workflow/references/cli-reference.md new file mode 100644 index 0000000..00f4234 --- /dev/null +++ b/skills/vlm-library-workflow/references/cli-reference.md @@ -0,0 +1,26 @@ +# VLM CLI Reference + +## Configuration +- `vlm config init`: Create default config. +- `vlm config show`: Display current settings. +- `vlm config validate`: Check config for errors. + +## Core Commands +- `vlm scan [--output PATH]`: Discover video files. +- `vlm parse [--input CSV] [--output JSON] [--inventory CSV]`: Parse filenames. +- `vlm enrich [--input JSON] [--output JSON] [--refresh-all]`: Fetch TMDB metadata. +- `vlm analyze [--input JSON] [--output JSON]`: Find gaps and duplicates. +- `vlm plan [--input JSON] [--analysis JSON] [--output JSON]`: Generate operations. +- `vlm execute [--plan JSON] [--confirm]`: Move/Rename files. +- `vlm rollback [--log PATH]`: Undo operations. + +## Management & Reporting +- `vlm quarantine list|add|restore`: Manage the `.quarantine` directory. +- `vlm report inventory|completeness|duplicates|summary`: Generate human-readable reports. +- `vlm state show|set|query|clear`: Track manual review status for files. + +## Important Config Options (`~/.vlm/config.yaml`) +- `library_root`: Path to the video collection. +- `templates`: Naming patterns for movies and series. +- `plan.duplicate_keep`: Strategy for duplicates (`by_quality`, `by_reputation`, `first_seen`, `manual`). +- `enrichment.api_keys`: TMDB and OpenAI keys. diff --git a/skills/vlm-library-workflow/references/command-recipes.md b/skills/vlm-library-workflow/references/command-recipes.md new file mode 100644 index 0000000..1864e80 --- /dev/null +++ b/skills/vlm-library-workflow/references/command-recipes.md @@ -0,0 +1,71 @@ +# VLM Command Recipes + +## Baseline + +Run from repository root unless user specifies otherwise. + +```bash +vlm --help +vlm config show +vlm config validate +``` + +If `vlm` is not on PATH, switch to: + +```bash +uv run vlm --help +uv run vlm config show +uv run vlm config validate +``` + +## End-to-End Pipeline + +```bash +vlm scan --output inventory.csv +vlm parse --input inventory.csv --output identities.json --inventory inventory.csv +vlm enrich +vlm analyze --input identities.json --output analysis.json --inventory inventory.csv +vlm plan --input identities.json --output plan.json --analysis analysis.json +vlm review-plan --input plan.json --output plan_manual_review.csv +vlm execute --plan plan.json +vlm execute --plan plan.json --confirm +``` + +## Artifact Expectations + +1. `inventory.csv`: discovered video files with filesystem and optional ffprobe metadata. +2. `identities.json`: parsed identities for movie/series/anime/other, optionally with embedded quality metadata. +3. `analysis.json`: completeness gaps and duplicate groups with quality comparison context. +4. `plan.json`: planned operations (`move`, `rename`, `quarantine`, `no-op`) and summary data. +5. `plan_manual_review.csv`: high-risk operations requiring human confirmation before `--confirm`. + +## Focused Workflows + +```bash +# Parse only, with metadata embedding +vlm parse --inventory inventory.csv + +# Analyze only +vlm analyze --input identities.json --output analysis.json --inventory inventory.csv + +# Plan from analysis-assisted duplicate decisions +vlm plan --input identities.json --analysis analysis.json --output plan.json + +# Roll back latest confirmed execution +vlm rollback +``` + +## Frequent Failure Triage + +1. Missing config: + Run `vlm config init`, then set `library_root` in `~/.vlm/config.yaml`. +2. Invalid config values: + Run `vlm config validate` and fix reported keys. +3. Missing input artifact: + Run the prerequisite stage (`scan` before `parse`, `parse` before `analyze`, `analyze` before analysis-driven `plan`). +4. Unexpected duplicate decisions: + Check `plan.duplicate_keep` in config and rerun `vlm plan --analysis analysis.json`. +5. Unsafe or undesired execution results: + Run `vlm rollback` and inspect plan before re-running `execute --confirm`. +6. `vlm` command not found: + Use `uv run vlm ...` fallback for the same subcommands. diff --git a/skills/vlm-library-workflow/references/dev-guide.md b/skills/vlm-library-workflow/references/dev-guide.md new file mode 100644 index 0000000..c8b6df2 --- /dev/null +++ b/skills/vlm-library-workflow/references/dev-guide.md @@ -0,0 +1,30 @@ +# VLM Developer Guide + +## Project Structure +- `src/vlm/cli.py`: Entry point and command definitions. +- `src/vlm/parser.py`: Regex-based filename parsing logic. +- `src/vlm/enrichment.py`: Pipeline for external metadata fetching. +- `src/vlm/providers/`: API implementations (e.g., TMDB). +- `src/vlm/planner.py`: Logic for mapping identities to filesystem operations. +- `src/vlm/executor.py`: Safe file manipulation and rollback logging. + +## Adding a New Command +1. Create a new module in `src/vlm/commands/`. +2. Define the command using `@click.command()`. +3. Register it in `src/vlm/cli.py` using `main.add_command()`. + +## Modifying the Parser +- The parser uses a sequence of regex patterns in `src/vlm/parser.py`. +- Add new patterns to the `PATTERNS` list or improve existing ones. +- Always run `pytest tests/test_parser.py` after changes. + +## Data Models +See `src/vlm/models.py` for core data structures: +- `VideoFile`: Basic file metadata. +- `MediaIdentity`: Parsed and enriched information. +- `PlanOperation`: Definition of a file move/rename/quarantine. + +## Testing +- **Unit Tests**: `pytest` +- **Property-based Tests**: `pytest tests/test_analysis_properties.py` (uses Hypothesis). +- **Integration Tests**: `pytest tests/test_reports_integration.py`. diff --git a/skills/vlm-library-workflow/references/workflow.md b/skills/vlm-library-workflow/references/workflow.md new file mode 100644 index 0000000..4a0152e --- /dev/null +++ b/skills/vlm-library-workflow/references/workflow.md @@ -0,0 +1,67 @@ +# VLM Workflow Guide + +This guide details the standard end-to-end process for organizing a video library using VLM. + +## 1. Setup and Discovery + +### Initialize Configuration +```bash +vlm config init +``` +Edit `~/.vlm/config.yaml` to set `library_root`. + +### Library Scan +```bash +vlm scan +``` +- **Goal**: Create `inventory.csv`. +- **Note**: Ensure `ffprobe` is installed for resolution and codec metadata. + +## 2. Identification + +### Filename Parsing +```bash +vlm parse --inventory inventory.csv +``` +- **Goal**: Create `identities.json`. +- **Why --inventory?**: It embeds video metadata (v2 schema) required for quality-based duplicate resolution. + +### Metadata Enrichment +```bash +vlm enrich +``` +- **Goal**: Update `identities.json` with TMDB data. +- **Troubleshooting**: If matches are missing, check `enrichment.api_keys` in config. + +## 3. Analysis and Planning + +### Detect Issues +```bash +vlm analyze +``` +- **Goal**: Create `analysis.json`. +- **Outputs**: Lists duplicate files and episode gaps in series. + +### Create Execution Plan +```bash +vlm plan --analysis analysis.json +``` +- **Goal**: Create `plan.json`. +- **Strategy**: VLM uses the `duplicate_keep` policy (default: `by_reputation`) to decide which files to keep and which to quarantine. + +## 4. Execution and Safety + +### Review the Plan +Open `plan.json` and check the `human_summary` field or the proposed `operations`. + +### Execute Changes +```bash +vlm execute # Dry-run +vlm execute --confirm # Actual operations +``` + +### Reverting Changes +```bash +vlm rollback +``` +Restores files using the latest log in `~/.vlm/rollback/`.