dedupe LLM_Evaluation zone: single-source tables, fix nav orphans, compress guidebook 2025
- Why Guide: remove duplicated 20-case/rubric/test-definition tables, link to authoritative Roadmap/Worksheet instead; 10-min action points to entries - Concept Map / What-Is: cross-link narrative vs quick-reference roles - Worksheet: annotate sections with authoritative sources - 04-Reference 01-04 <-> archive/01: bidirectional resource links - archive/00-Material-List: shrink guidebook listing, now reachable from READMEs - guidebook: compress 06-Yearly-Dives 2025 section (-> 08-2025-Edition §3), add old/new version nav banners to 01-05, reverse links in 08 - README/Start-Here: link Learning Board (was orphaned)
This commit is contained in:
+4
@@ -118,6 +118,8 @@ article.mdx(组装层,含正文过渡小节、saturation/contamination 定
|
||||
|
||||
## §3 2025 评测全景(2025-evaluations-for-useful-models + article.mdx)
|
||||
|
||||
> 旧版对应:[[06-Yearly-Dives]] 的 2025 节(旧版压缩摘要,含旧版独有的"核心论点")。
|
||||
|
||||
> 新版先给出两个贯穿全篇的核心概念(来自 article.mdx 的 "Evaluating with existing benchmarks" 引言):
|
||||
>
|
||||
> - **Saturation(饱和)**:模型在 benchmark 上的表现超过人类表现。更广义地指数据集失去模型间区分力、不再有用——"如果所有模型分数都接近最高分,它就不再是 discriminative benchmark,就像拿学前班题目考高中生:成功说明不了什么(虽然失败能说明问题)"。
|
||||
@@ -273,6 +275,8 @@ article.mdx(组装层,含正文过渡小节、saturation/contamination 定
|
||||
|
||||
## §5 设计自动评测(designing-your-automatic-evaluation + article.mdx)
|
||||
|
||||
> 旧版对应:[[01-Automatic-Benchmarks]] §3(设计自动评测)与 §4(常用评测数据集盘点,新版不再渲染正文)。
|
||||
|
||||
新版把"设计自动评测"重写为完整方法论。**与旧版不同的章节结构**(原文实际顺序):
|
||||
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user