reconcile LLM_Evaluation zone: align stage numbering and scopes, standardize full-path wikilinks, slim old guidebook notes, dedupe templates
- fix stage-numbering conflict (README vs 05-Progress) and unify stage-1 reading scope - resolve AWS workshop prerequisite contradiction in 04-Reference/01 - convert medium-path wikilinks to vault-root paths (~30 links), fix .pyy typos, annotate ragas fork, unify archive status, add 01-/02- README hubs - compress old-version guidebook notes (01, 05) into pointers; add 2026 reading guidance to 00-Overview - dedupe project templates and remove embedded template copy in 03-Practice/README
This commit is contained in:
@@ -335,7 +335,7 @@ Run 记录:这个版本的系统在某个时间、某个配置下实际输出
|
||||
| 暂时不做 | 为什么现在不做 | 什么时候再学 |
|
||||
|---|---|---|
|
||||
| 训练 / 微调模型 | 你还没有稳定的质量标准,不知道该用什么数据改进 | 能稳定设计 case、rubric 与回归集之后。 |
|
||||
| LLM-as-a-Judge | 自动评分会掩盖 rubric 是否清楚 | 手工盲评至少 20—50 条并复盘分歧之后。 |
|
||||
| LLM-as-a-Judge | 自动评分会掩盖 rubric 是否清楚 | 手工盲评至少 20—50 条并复盘分歧之后(正式校准用 50–100 条代表样本,见路线图第七节)。 |
|
||||
| CI 门禁、平台和仪表盘 | 它们放大已有流程,不会创造流程 | 手工重跑开始重复、容易漏步骤之后。 |
|
||||
| 大规模红队 | 范围广、风险分类复杂 | 有一个具体系统边界,如 RAG 注入或工具越权之后。 |
|
||||
| 1000 条数据 | 数量会掩盖设计问题 | 你能明确说出每个类别为何存在之后。 |
|
||||
|
||||
Reference in New Issue
Block a user